0% found this document useful (0 votes)
7 views192 pages

Module 2

The document discusses descriptive statistics, focusing on data collection, population definitions, and the importance of sampling for statistical inference. It explains different measurement scales (nominal, ordinal, interval, and ratio) and their implications for data analysis, as well as the differences between induction and deduction in statistical reasoning. Additionally, it covers univariate analysis, types of frequencies, and the role of descriptive statistics in summarizing and visualizing data.

Uploaded by

Kunal MK
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views192 pages

Module 2

The document discusses descriptive statistics, focusing on data collection, population definitions, and the importance of sampling for statistical inference. It explains different measurement scales (nominal, ordinal, interval, and ratio) and their implications for data analysis, as well as the differences between induction and deduction in statistical reasoning. Additionally, it covers univariate analysis, types of frequencies, and the role of descriptive statistics in summarizing and visualizing data.

Uploaded by

Kunal MK
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Descriptive Statistics

Introduction
• Historical Context & Data Collection:

– Caesar Augustus’ census is an early example of data collection at a large scale.

– It ensured a complete dataset of the Roman Empire’s population, just like


modern businesses and governments conduct surveys or track user data.

• Definition of a Population:
A population can refer to any group being studied, such as:
– People in a country.
– Employees in an organization.
– Animals in a zoo.
– Cars of an institution.
– Research institutions in a country.
– Products, such as nails from a machine.
Challenges in Surveying Populations:
• Practical Limitations:
– In some cases, it is impossible to include the
entire population (e.g., collecting all nails ever
produced by a machine).
• Cost Constraints:
– Surveying all members of a population can be
cost-prohibitive (e.g., surveying all citizens in an
election).
• Exceptions exist for small populations where a
complete survey may be feasible.
Sampling
• Sampling is the process of selecting a subset
of individuals, items, or observations from a
larger population to analyze and make
inferences about the whole population.

• It is commonly used in statistics, research, and


data analytics when studying an entire
population is impractical or too costly.
Importance of Sampling

• Sampling is a practical way to study populations


when surveying the entire population is not
feasible.

• By analyzing a subset (sample), it is possible to


estimate specific values for the entire population.

Example of Sampling

• A sample can be used to estimate values such as


the proportion of votes for a political party.
Statistical Inference
• Definition: Generalizing knowledge from a sample to the entire
population is called statistical inference (or induction).

• Key Idea: Statistical inference involves making conclusions about a


population based on information from a sample.

Variability in Sampling

• Different samples from the same population can yield different results.

• The value inferred for the population may vary depending on the
sample chosen.
Uniqueness of Population Values

• The true value for the entire population is


unique and fixed, but it can only be obtained if
the entire population is surveyed.

Sample Size and Accuracy

• Larger samples result in estimates that are


closer to the actual value for the population.
Difference Between Induction and Deduction
• Induction:

– Generalizes from a sample to the population.

– Example: Estimating a population's characteristics based


on a sample.

• Deduction:

– Particularizes from the population to a sample.

– Example: Predicting the characteristics of a sample based


on known population data.
Example of Deduction
• Problem:
– Given the population of a university, what is the probability of
selecting individuals from two different continents in a random
sample of size 10?

– This involves deducing sample characteristics from population data.

Probabilities and Deduction

• Probability problems typically fall under deduction.

• The goal is to use known information about the population to


infer the likelihood of specific outcomes in a sample.
• In data analytics projects, data samples are often
retrieved using tools like SQL.

• The goal is to extract insights from the data.

Challenges with Large Data

• Data samples can be very large (hundreds or


thousands of instances).

• Directly analyzing such large datasets can be


overwhelming and may lead to confusion rather
than understanding.
Role of Descriptive Statistics

• Definition:

– Descriptive statistics focuses on methods to summarize and


visualize data samples.

• Purpose:

– Simplifies large datasets into manageable and understandable


formats.

– Provides insights through summarization (e.g., averages,


medians, ranges) and visualization (e.g., graphs, charts).
A figure is mentioned to depict relationships between the discussed concepts (e.g.,
induction, deduction, descriptive statistics, and data samples).
Categorization of Data Analysis
• Univariate Analysis

• Bivariate Analysis

• Multivariate Analysis
Scale Types in Statistics
• Scale types refer to the different ways data can be
measured and categorized.

• The provided dataset includes different types of


variables, which can be classified into two main
categories:
– Qualitative:
• Nominal Scale (No Order, Just Labels)
• Ordinal Scale (Ordered but No Fixed Differences)

– Quantitative:
• Interval Scale (Ordered, Equal Intervals, No True Zero)
• Ratio Scale (Ordered, Equal Intervals, True Zero)
• The amount of information we can extract from
data depends on the scale type used to express
it.
• The four main types of measurement scales,
arranged from least to most informative, are:

• Nominal Scale (Least Informative)


– Purpose: Categorization
– Examples: Colors (red, blue, green), Gender (male,
female), Types of birds (chicken, duck, turkey)
– Key Feature: No meaningful order or numerical
interpretation
– Limitations: Cannot compare values or perform
mathematical operations
Ordinal Scale
– Purpose: Ranking
– Examples: Customer satisfaction (poor, average,
good, excellent), Race positions (1st, 2nd, 3rd)
– Key Feature: Order matters, but differences
between values are not meaningful
– Limitations: We know the ranking but not the
exact difference between ranks
Relative Scale (Ratio Scale)

– Purpose: Measurement with meaningful ratios

– Examples: Height, Weight, Speed, Kelvin temperature

– Key Feature: Has a true zero, allows meaningful ratios (e.g., 10 kg is twice 5 kg)

– Limitations: Still depends on units of measurement

Absolute Scale (Most Informative)

– Purpose: Exact measurement with a natural unit

– Examples: Number of eggs laid by a hen, Number of students in a class

– Key Feature: True zero and fixed unit of measurement (e.g., 5 eggs means exactly 5
eggs, independent of any scale)

– Limitations: Rarely used in all scientific measurements (since many measurements


rely on units like kg, meters, etc.)
Difference Between Ordinal and Nominal Scales

Feature Ordinal Scale Nominal Scale


Definition A scale where categories have A scale where categories are
a meaningful order or ranking. just labels without any order.
Order Exists? Yes (Ranked) No (No Ranking)
Equal Differences? No (Differences between ranks No (No numerical meaning)
are not measurable)
Mathematical Operations Can compare which is Only used for categorization,
higher/lower, but can’t no mathematical comparisons.
measure exact differences
Examples Company Rating (Good, Bad) Gender (Male, Female) → No
→ "Good" is better than "Bad" ranking, just categories.
but the gap is undefined.
Real-Life Example Education Level (High School, Blood Type (A, B, AB, O) → No
Bachelor's, Master's, PhD) → type is greater than the other.
There is a progression.
Difference Between Absolute and Relative Scales
Scale Type Definition Key Feature Example
Absolute Scale A scale where Has an absolute Weight (kg), Height
measurements have zero, allowing ratio (cm), Age (years),
a true zero point, comparisons (e.g., Income ($) → 0 kg
meaning zero twice as much). means no weight.
represents the
complete absence
of the measured
property.
Relative Scale A scale where zero No true zero, so Temperature (°C or
is arbitrary and ratios are not °F), IQ scores, pH
does not indicate a meaningful. levels → 0°C does
total absence of the not mean "no
property. temperature."
Scale Type Reason Example
Variable
Values
Andrew,
No ranking or order, just
Name of Contact Nominal Carolina,
labels
James
Ordered categories (Good >
Company Rating Ordinal Bad), but differences are not Good, Bad
measurable
Attribute Measurement Zero Meaning Ratios Make Example
Scale Sense?

Height Ratio No height Yes 180 cm is twice 90 cm


Weight Ratio No weight Yes 100 kg is twice 50 kg
Temperature Interval Not absence of No 20°C is not twice 10°C
(°C, °F) temperature

Temperature Ratio Absolute zero (no Yes 200K is twice 100K


(Kelvin) thermal energy)
Nominal → Ordinal → Relative (Ratio) → Absolute
(Least Informative) → (Most Informative)
Scale Type Allowed Operations Example
Nominal Equality (=), Inequality (≠) Checking if two car brands are the same
(Toyota = Honda?)
Ordinal All Nominal operations + Ranking students in a class (Is Alice ranked
Ordering: Greater than (>), higher than Bob?)
Less than (<), Greater than or
equal to (≥), Less than or
equal to (≤)
Relative All Ordinal operations + Weight difference (How much heavier is a
(Ratio) Addition (+), Subtraction (−) 100 kg person than a 50 kg person?)
Absolute All Relative operations + Counting objects (If one farm has 10 cows
Multiplication (×), Division (÷) and another has 5, one has 2× more)
• This all means that when we have data expressed on
an absolute scale we can convert it to any of the
other scales.

• When we have data expressed on a relative scale we


can convert it in any scale of the two qualitative scale
types.

• When we have data expressed on an ordinal scale we


can express it in a nominal scale.

• But we should be aware that converting a more


informative scale into a less informative one involves
loss of information.
This example demonstrates how the same weight data (kg) can be expressed in
different measurement scales, each with different levels of information.

1. Absolute → Relative Scale


• Process: We modify the scale by subtracting 10 kg from
all values.
• Effect:
– The old zero (0 kg) becomes -10 on the new scale.
– The new zero is the old 10 kg.
– Why is 80 kg no longer twice 40 kg?
• On the original absolute scale: 80 kg / 40 kg = 2 (makes sense).
• On the new relative scale:
– Old 80 kg → New 70 kg (80 - 10)
– Old 40 kg → New 30 kg (40 - 10)
– 70 kg / 30 kg ≠ 2, so the ratio property is lost.
• Key Insight: Since the transformation changes the reference point,
it no longer makes sense to say one weight is "twice" another.
2. Relative → Ordinal Scale
• Process: We create categories based on weight
ranges.
• Effect: Instead of exact weight values, we classify
people as:
– Fat if weight > 80 kg
– Normal if 65 kg < weight ≤ 80 kg
– Thin if weight ≤ 65 kg
• What changed?
– We now only know who is heavier but not by how much.
– We lose precise numerical values.
– The choice of 65 kg and 80 kg is arbitrary, but it follows a
rationale
3. Ordinal → Nominal Scale
• Process: We replace ordered labels (Fat,
Normal, Thin) with arbitrary symbols (B, A, C).

• Effect:
– The new categories B, A, C do not have an
inherent order.
– We lose ranking information—we cannot tell if B
is heavier than A.
– This is the least informative scale (only
classification, no ranking or numeric values).
Understanding Scale Types vs. Data Types in Software
• In software, when working with data, we must choose data types for each
attribute.
• However, data types and scale types are not the same, even though they
are related.
1. Scale Types vs. Data Types

Concept Definition Example


Scale Type A way to classify data based on how Nominal, Ordinal,
it can be measured and interpreted. Relative (Ratio), Absolute
Data Type The format in which data is stored Integer, Float, Text, Date
in a software package.

•Scale types describe the level of measurement (e.g., categorical vs.


numerical).

•Data types define how the data is stored in memory and the operations
allowed on it.
Common Data Types in Software
Data Type Description Scale Type Example

Text Stores text (e.g., names, categories) Nominal (e.g., "Toyota", "Honda")

Character Stores a single letter or symbol Nominal (e.g., Gender: "M", "F")

Factor Stores categorical data with predefined levels Ordinal (e.g., "Low", "Medium", "High")

Integer Stores whole numbers Absolute or Relative (e.g., "5 apples",


"Age in years")
Real/Float Relative (e.g., weight in kg, temperature
Stores decimal numbers (continuous values) in °C)
Timestamp Stores date and time Absolute (e.g., "2024-03-11 10:00 AM")

Date Stores only date information Absolute (e.g., "March 11, 2025")

Discrete vs. Continuous Numeric Data


Discrete values (e.g., integer) → Countable, fixed values (e.g., number of eggs laid).

Continuous values (e.g., float, real) → Can take any value within a range (e.g., temperature,
weight).
Numbers Do Not Always Mean Quantitative Data
• Just because an attribute is expressed as a number does not
mean it is quantitative (i.e., a relative or absolute scale).
• It could still be nominal or ordinal, depending on its purpose.
• Example: Numeric Codes vs. Quantitative Values
Example Scale Type Reason
ID Card Number (e.g., 12345) The number is just a label; it does not have
Nominal
mathematical meaning.
Postal Code (e.g., 560001, 90210) Nominal These numbers represent locations but cannot be
added, subtracted, or ordered meaningfully.
Rank in a Competition (e.g., 1st, Ordinal The ranking shows order, but differences between
2nd, 3rd) ranks are not necessarily equal.
Temperature in °C (e.g., 30°C, Relative Differences matter, but 0°C does not mean "no
40°C) temperature."
Weight in kg (e.g., 50kg, 100kg) Absolute The numbers have a true zero and can be used for
multiplication and division.
Descriptive Univariate Analysis
• Descriptive Univariate analysis focuses on
summarizing and interpreting a single variable
in a dataset.

• It helps understand
– the distribution,
– central tendency, and
– variability of the data
Univariate Frequencies
• Univariate frequency analysis helps
summarize how often each value of a single
variable appears in a dataset.

• This is particularly useful for understanding


categorical or discrete numerical data
distributions.
Types of Frequencies
• Absolute Frequency
– This is the raw count of how many times a
particular value appears in the dataset.
– Example: If "Male" appears 30 times in a dataset
of 50 individuals, its absolute frequency is 30.

• Relative Frequency
– This is the proportion or percentage of
occurrences of a specific value compared to the
total number of observations.
Cumulative Frequency
• This represents the sum of frequencies up to a certain
value.

• It helps understand data distribution.

• Often used in numerical data, where values are ordered.

• Example: If test scores are categorized as follows:

– 0-50: 5 students
– 51-70: 10 students
– 71-90: 15 students
– 91-100: 20 students

• The cumulative frequency for the "71-90" range is 5 + 10 +


15 = 30 students.
Absolute and Relative Frequency for "Company“
• The "Company" attribute in Table 2.1 has two
categories: Good and Bad.
• The absolute and relative frequencies are:

• Absolute frequency counts how many times each


category appears.
• Relative frequency shows the percentage of
occurrences.
Absolute, Relative, and Cumulative Frequency for
"Height“
• The height values from Table 2.1 are:
175, 195, 172, 180, 168, 173, 180, 165, 158, 163,
190, 172, 185, 192

• Sorted height values in ascending order:


158, 163, 165, 168, 172, 172, 173, 175, 180, 180,
185, 190, 192, 195
•Cumulative frequency shows the number of occurrences ≤ a given value.
•Relative cumulative frequency represents the
percentage of total occurrences ≤ a given value.
• Categorical data ("Company"): The frequency
table is easy to interpret since there are only two
categories.
• Numerical data ("Height"): Since values are
spread out, frequency tables can become less
useful.
• Empirical frequency distribution: The column
"Relative Frequency" defines the frequency of
each value.
• Empirical cumulative distribution function: The
column "Relative Cumulative Frequency" helps
analyze percentiles and data spread.
Probability Mass Function (PMF) for Discrete
Data
• Used for discrete attributes (e.g., integer
values like counts, number of eggs laid by a
hen, etc.).
• Represents probabilities of specific values
occurring.
• The total sum of probabilities must equal 1
(i.e., 100%).
• Example: If the probability of rolling a 3 on a
fair die is 1/6, then P(X=3)=1/6
Probability Density Function (PDF) for
Continuous Data
• Used for continuous attributes (e.g., height,
temperature, weight).
• The probability of any exact value is theoretically zero
because continuous values can take infinite decimal
places.
• Instead of exact probabilities, PDFs describe relative
densities over a range.
• The area under the PDF curve equals 1, representing
the total probability.
• Example:
• The probability that a person’s height is exactly 175 cm
is zero, but we can find the probability that it falls
within a small range (e.g., 175–176 cm).
• Why Probability Density is Needed for
Continuous Data
• Imagine measuring height with infinite
precision (e.g., 175.2384761… cm).
• The chance that two people have the exact
same height is nearly zero.
• Instead, we use density: how likely a value
falls within an interval.
• The probability of height between 175 cm and
176 cm is a nonzero value, represented by the
area under the curve.
Univariate Data Visualization
Pie Chart
• A pie chart is mainly used for nominal (categorical) data,
where categories have no inherent order (e.g., "Company" in
your dataset with values Good and Bad).

Why Pie Charts Work for Nominal Data:


• They visually show proportions of different categories.
• The total percentage sums to 100%.
• It is not recommended for ordinal or quantitative data since
those have an order, and other charts like bar charts or
histograms are better.

Example: Pie Chart for "Company" Data


• Good: 7 people (50%)
• Bad: 7 people (50%)
• A pie chart would show these two categories as equal slices.
Bar Chart
• Bar charts are typically used for qualitative
(categorical) data, but they can also be applied to
quantitative data in specific cases.

Key Points About Bar Charts:


• Better Comparisons:
– Easier to compare sizes of bars than pie chart slices.
• Order Matters:
– If categories have a natural order (ordinal data), bars
should be arranged in increasing order.
• Use with Limited Quantitative Values:
– Example: Rolling a die (1-6 outcomes) → Counts can be
represented as bars.
– Example: Exam scores (0-20 scale) → Bars for each score
value.
When to Use a Bar Chart Instead of a Pie Chart
• If you have many categories, a pie chart
becomes hard to read.
• If your data has an inherent order, bar charts
preserve that order.
• When the differences between categories
matter, bar charts make comparisons easier.
Line Charts: Best for Time-Series Data
• Line charts are particularly useful for quantitative
data where the horizontal axis represents time or
another continuous variable.

Key Characteristics of Line Charts:


• Used for time series data
– Example: Daily temperature readings in a town.
– Example: Stock market trends over time.
• Data points are connected by lines, making it easier
to spot trends and patterns.
• Equally spaced observations on the horizontal axis
(e.g., days, months, years).
Common Use Cases:
• Tracking changes over time (e.g., temperature,
sales, stock prices).
• Comparing trends across different groups (e.g.,
GDP growth of multiple countries).
• Identifying patterns like seasonality or cyclic
behavior.
Why Not Use Line Charts for Categorical Data?
• Line charts assume a natural numerical order, so
they should not be used for qualitative
(categorical) data like "Company" or "Gender".
• Instead, bar charts or pie charts are better for
comparing categories.
Area Charts: Best for Comparing Time Series &
Distributions
Key Characteristics of Area Charts:
• Used for Time Series & Distribution Functions
– Helps visualize how values change over time or across a distribution.

• Similar to Line Charts, But Filled Below the Line


– The area under the curve is shaded to highlight magnitude.

• Great for Comparing Multiple Data Series


– Example: Comparing temperature trends in different cities over time.
– Example: Comparing stock market growth of multiple companies.

• Can Show Probability Density Functions (PDFs)


– Used in statistics to understand data concentration (e.g., heights of
people).
Use Cases:
• Time-series analysis (e.g., rainfall over
months).
• Comparing distributions (e.g., income
distribution across regions).
• Showing cumulative data (e.g., total sales
over months).
Histograms: Understanding Data
Distribution
• Histograms are essential tools for visualizing
empirical distributions of quantitative attributes.

• Unlike bar charts (representing categorical data),


histograms help us see patterns, trends, and
variations in numerical data by grouping values
into intervals, bins, or cells.
Key Features of Histograms
• Used for Quantitative (Numerical) Data :
– Ideal for continuous attributes like height, weight,
temperature, income, and exam scores.
– Example: A histogram can show how students’ exam
scores are distributed (e.g., 0-10, 11-20, 21-30, etc.).
• Bins (Intervals) Group Data for Better Insights:
– Since continuous data can have many unique values,
histograms reduce sparsity by aggregating values into
intervals.
– Example: The heights of people could be grouped into
[150-160], [160-170], and [170-180] cm rather than
showing every individual height.
• More Informative Than Bar Charts:
– Unlike bar charts (which are for categorical data),
histograms show frequency distributions effectively.
– Bar charts have spaces between bars, while
histograms do not, ensuring continuity in data
representation.
• No Gaps Between Bars:
– Since histograms represent numerical values that
flow continuously, bars are placed adjacent to each
other.
– The lack of gaps reinforces the idea of continuity in
the data.
Choosing the Right Number of Bins (Cells)
• A critical decision in histogram design is how many bins to use.
• The number of bins affects how well the data is represented:

Too Few Bins → Data looks too generalized, hiding patterns.


Too Many Bins → Data looks too noisy, making patterns harder to interpret.

👉 Rule of Thumb: The number of bins is often chosen as √(number of


values).

• Example: If there are 100 values, a good starting point is 10 bins.

• The best choice depends on the dataset and the specific analysis needs.

✔️Equal-sized bins are typically used for simplicity.


✔️If varying bin widths are used, the height of the bars should be adjusted
proportionally to maintain correct representation.
Empirical vs. Probability Distributions
• Histograms help us understand empirical distributions
based on a sample dataset, but they are different from
theoretical probability distributions, which describe the
behavior of an entire population.
1. Empirical Distribution
– Created from real-world sample data.
– Used to study how data points are distributed.
– Example: A histogram of people’s heights in a particular city or
school.
2. Probability Distributions
– Based on mathematical models and theoretical expectations.
– Used to predict the probabilities of different outcomes in a
population.
– Example: The normal distribution (bell curve) is often used to
model height distributions in an entire country.
Cumulative Distributions: Seeing Data Trends Clearly
• A cumulative distribution function (CDF) represents the
percentage of data points that are less than or equal to a
certain value.
• The empirical cumulative distribution function (ECDF) is
derived from a sample.
Key Insights from CDFs:
• Flat (horizontal) sections → Indicate values that appear less
frequently.
• Steep (vertical) sections → Indicate values that appear more
frequently.
• Step-wise structure → Comes from sampling and
measurement precision limits (e.g., rounding heights to the
nearest cm).
• Example of Cumulative Distribution Interpretation:
If a height cumulative distribution function shows a steep
increase between 160-170 cm, it means many people fall
within that height range.
Stacked Bar Charts for Analyzing Multiple Attributes

• While histograms show single-variable distributions, stacked bar


charts allow us to analyze the relationship between two
attributes.

Example: Stacked Bar Chart for ‘Company’ & ‘Gender’

• Instead of just showing the frequency of ‘Good’ vs. ‘Bad’


companies, we can further divide the bars into Male (M) and
Female (F) categories.

• This helps us see if one gender is more associated with ‘Good’


or ‘Bad’ companies.
Univariate Statistics
• A statistic is a numerical descriptor that
summarizes a characteristic of a dataset,
either for a sample or a population.

• These statistics help in understanding the


central tendency and variability of the data.

• There are two main groups of univariate


statistics:
– location statistics
– dispersion statistics.
Location univariate statistics: Identifying Key Values in a Dataset
• Location statistics are used to pinpoint specific values in a dataset based on
their position.
• These statistics help summarize the central and extreme values in a dataset.
Key Location Statistics:
• Minimum: The lowest value in the dataset.
• Maximum: The highest value in the dataset.
• Mean (Average): The sum of all values divided by the number of values.
• Mode: The most frequently occurring value in the dataset.
• First Quartile (Q1 - 25th percentile): The value that is larger than 25% of the
data.
• Median (Second Quartile, Q2 - 50th percentile): The middle value that splits the
dataset into two equal parts.
• Third Quartile (Q3 - 75th percentile): The value that is larger than 75% of the
data.
Measures of Central Tendency
• Three important statistics help define the "center"
of a dataset:
– Mean (Average) – Sensitive to extreme values.
– Median – More robust to outliers.
– Mode – Useful for categorical data or detecting the
most common value.
• The mean (or average), median, and mode are
known as measures of central tendency, because
they return a central value from a set of values
• These statistics are crucial for understanding the
distribution of data
Example: the attribute “weight” from our data set
• Extract the weight values from the table and compute
the following location statistics:
• Minimum
• Maximum
• Mean (Average)
• Mode (most frequent value)
• First Quartile (Q1, 25th percentile)
• Median (Q2, 50th percentile)
• Third Quartile (Q3, 75th percentile)
1. Mean (Average)
The mean is calculated as:
Mean=∑X / N​
where:
• ∑X is the sum of all values
• N is the number of values

Dataset values for Weight (kg):


77, 110, 70, 85, 65, 75, 75, 63, 55, 66, 95, 72, 83, 115
Step 1: Compute the sum of all values
77+110+70+85+65+75+75+63+55+66+95+72+83+115=1106

Step 2: Divide by the number of values (N = 14)


Mean=1106/14=79.0
2. Mode (Most Frequent Value)
• The mode is the most frequently occurring value
in the dataset.
• Dataset values (sorted):
55, 63, 65, 66, 70, 72, 75, 75, 77, 83, 85, 95, 110,
115
• The value 75 appears twice, while all others
appear once.
• Mode = 75 kg
3. First Quartile (Q1, 25th percentile)
• Quartiles split the data into four equal parts.
Step 1: Sort the dataset
Sorted values:
• 55,63,65,66,70,72,75,75,77,83,85,95,110,115
4. Median (Q2, 50th percentile)
• The Median (Q2) is the 50th percentile or the
middle value in the sorted dataset.
• Step 1: Find the middle index
5. Third Quartile (Q3, 75th percentile)
• The Third Quartile (Q3) is the 75th percentile,
calculated as:
Box Plot
• A box plot (or box-and-whisker plot) is a great way
to visualize the spread and central tendency of a
dataset.
It displays five key statistics:
• Minimum – The smallest value in the dataset
• First Quartile (Q1, 25th percentile) – The median of
the lower half of data
• Median (Q2, 50th percentile) – The middle value of
the dataset
• Third Quartile (Q3, 75th percentile) – The median
of the upper half of data
• Maximum – The largest value in the dataset
• The box represents the inter quartile range (IQR
= Q3 - Q1), showing the middle 50% of the data.
• The horizontal line inside the box represents the
median (Q2, 50th percentile).
• The whiskers extend to the minimum and
maximum values (unless there are outliers).
• Outliers, if present, are plotted as individual
points beyond the whiskers.
• The bottom and top points of the plot are
respectively the minimum and the maximum.
• The bottom and the top of each box are respectively
the first and the third quartiles.
• The horizontal line in the middle of the box is the
median.
• The closer each of the points is, the more frequent
the values between these points are.
• for example, the distance between the first quartile
and the median.
• Twenty-five percent of the values are found
between these two statistics.
• Mean, median and mode are measures of central
tendency
When to Use Mean, Median, and Mode?
Mean (Average)
• Use When:
– Data is quantitative (numerical).
– Data is symmetrically distributed (no extreme
outliers).
– You want a measure that considers all values.
• Avoid When:
– The data has outliers (extremely high or low values)
because they can skew the mean.
– The data is skewed (not symmetric).
• Example:
– Average height of students in a class.
– Average income of employees in a company (but be
careful of very high salaries that can skew the mean).
Median (Middle Value)
• Use When:
– Data is quantitative or ordinal (numerical or ranked).
– Data has outliers or is skewed.
– You want a measure that represents the middle of the
dataset.
• Avoid When:
– You need to include all values in the calculation.
• Example:
– Median house price (useful because real estate has
extreme price variations).
– Median income (more representative than mean
income due to high-income outliers).
Mode (Most Frequent Value)
• Use When:
– Data is nominal (categorical), ordinal, or quantitative.
– You want to find the most common value.
– The dataset has repeated values.
• Avoid When:
– Every value appears the same number of times (no
mode).
– The dataset is too small to determine meaningful
frequency.
• Example:
– Most popular car color in a city (nominal).
– Most common exam score in a class (quantitative).
Dispersion Univariate statistics
• A dispersion statistic measures how spread
out or distant different values are in a dataset.

• It provides insights into variability and


consistency in the data.
The most common dispersion statistics are:

• amplitude: the difference between the maximum


and the minimum values
• Inter quartile range: is the difference between
the values of the third and first quartiles
• mean absolute deviation: a measure for the
mean absolute distance between the
observations and the mean. Its mathematical
formula for the population is:
• where n is the number of observations and 𝜇x
is the mean value of the population.
• In this case the distance between an
observation and the mean contributes to the
MAD in linear proportion to that distance.
• For example, one observation with a distance
to the mean of 4 will increase the MAD by the
same amount as two observations each with a
distance to the mean of 2.
• standard deviation: another measure for the typical
distance between the observations and their mean. Its
mathematical formula for the population is:

• where n is the number of observations and 𝜇x is the mean


value of the population.
• In this case the distance between an observation and the
mean contributes to the standard deviation in quadratic
proportion to that distance.
• For example, one observation with a distance to the mean
of 4 will increase 𝜎 more than two observations each with
a distance to the mean of 2.
• The square of the sample deviation is termed the variance
and is denoted as 𝜎2.
• It measures how spread out the population values are
around the mean.
• All of these dispersion statistics are only valid
for quantitative scales.
Calculate the amplitude, interquartile range (IQR), mean absolute deviation (MAD),
and standard deviation (σ) for the weight attribute from the company dataset.

Step 1: Extract Weight Data


Weights (kg) from the dataset:
77,110,70,85,65,75,75,63,55,66,95,72,83,115
Step 3: Calculate Interquartile Range (IQR)
• IQR is the difference between the third quartile (Q3​) and the first quartile
(Q1).
Compute the Mean Absolute Deviation (MAD) for the weight attribute step by step.

Step 2: Compute Absolute Deviations


for each weight value xi, compute ∣xi−μ∣:
55-79=24 85-79=6
63-79=16 95-79=16
65-79=14 110-79=31
66-79=13 115-79=36
70-79=9
72-79=7
75-79=4
77-79=2
83-79=4
compute the standard deviation (s) for the weight attribute
• When working with a sample rather than the
entire population, we adjust our calculations
using n - 1 instead of n.

• This adjustment, known as Bessel’s correction,


helps correct bias in estimating the population
variance and standard deviation.
Why Do We Use n−1 Instead of n?
• If we use n, we underestimate the population
variance.
• Using n−1 corrects this bias, ensuring an
unbiased estimator of variance.
Common Univariate Probability Distributions

• Each attribute has its own probability distribution.


• Many common attributes follow functions for which the distribution is
already known.
• Describes how the values of a single random variable are distributed.
Discrete Distributions
These are used when the variable can take only specific, separate values.
• Bernoulli Distribution: Models a binary outcome (success/failure, 1/0) with
probability p.
• Binomial Distribution: Describes the number of successes in n independent
Bernoulli trials.
• Poisson Distribution: Models the number of events occurring in a fixed
interval of time or space when events occur independently at a constant
rate.
• Geometric Distribution: Describes the number of trials needed to get the
first success in a sequence of Bernoulli trials.
Continuous Distributions
These apply to variables that can take on any value within a range.
• Uniform Distribution: All values within an interval [a,b] are equally
likely.
• Normal (Gaussian) Distribution: Bell-shaped curve characterized by
mean μ and standard deviation σ. Many natural processes follow this
distribution.
• Exponential Distribution: Models the time between events in a
Poisson process.
• Gamma Distribution: Generalization of the exponential distribution,
often used in reliability analysis.
• Beta Distribution: Defined on [0,1], useful in Bayesian statistics and
probability modeling.
• Log-Normal Distribution: A variable whose logarithm is normally
distributed
• Two of these distributions:
– the uniform
– the normal(Gaussian).
• Both are continuous distributions and have
known probability density functions.
Uniform Distribution
• The uniform distribution is a very simple distribution.
• The frequency of occurrence of the values is uniformly distributed in a
given interval of values.
• An attribute x that follows a uniform distribution with parameters a and b,
respectively the minimum and maximum values of the interval.
• The probability of x < 0.3 is given by the proportion of the
area taken by this term, as shown in Figure
The Normal (Gaussian) Distribution

• The Normal Distribution, also called the Gaussian Distribution, is widely


used in statistics because of its natural occurrence in many real-world
phenomena. It is denoted as:
Connection to the Central Limit Theorem (CLT)
• The Central Limit Theorem (CLT) states that the sum (or
mean) of many independent random variables, regardless of
their original distribution, tends to follow a normal
distribution as the sample size increases.
• This is why the normal distribution is commonly observed in
nature and used in statistical inference.
Example Applications
• Natural Phenomena: Heights of people, IQ scores,
measurement errors.
• Finance: Stock returns, risk modeling.
• Machine Learning & Data Science: Feature distributions,
probabilistic modeling.
The normal distribution is a symmetric and continuous distribution, as shown in Figure
Descriptive Bivariate Analysis

• Bivariate analysis examines the relationship between


two variables.
• The methods used depend on the scale type of the
variables:
– Two Quantitative Variables → Use scatter plots, correlation
coefficients (e.g., Pearson, Spearman), and regression
analysis.
– One Quantitative & One Qualitative (Nominal/Ordinal)
Variable → Use box plots to compare distributions across
categories.
– Two Qualitative Variables → Use contingency tables and chi-
square tests to analyze associations.
Quantitative vs. Quantitative (Numerical Data)
• Scatter Plots: Visualize relationships and
patterns between two numeric attributes.
• Correlation Coefficients:
– Pearson’s r: Measures linear correlation.
– Spearman’s ρ: Measures rank-based correlation
(useful for non-linear relationships).
• Regression Analysis: Models relationships
(e.g., simple linear regression: y=mx+b).
Quantitative vs. Qualitative (Categorical [Link])
• Box Plots: Show distribution of a numeric variable
across different categories.

• ANOVA (Analysis of Variance): Tests differences in


means between groups.

• T-tests: Compare means between two groups.


Qualitative vs. Qualitative (Categorical Data)
• Contingency Tables: Show frequency distribution
of two categorical variables.

• Chi-Square Test: Tests for independence between


categorical variables.

• Bar Charts & Mosaic Plots: Visualize categorical


relationships.
Two Quantitative Attributes

• In a data set whose objects have n attributes, each object can be


represented in a n-dimensional space:
– a space with n axes, each axis representing one of the attributes.
• The position occupied by an object is given by the value of its
attributes.
• There are several visualization techniques that can visually show the
distribution of points with two quantitative attributes.
• One of these techniques is an extension of the histogram called a
three-dimensional histogram.
• 3D Histogram (or Binned Scatter Plot)
• A histogram extended to two variables, where data points are
grouped into bins in a 3D space.
• Each bin’s height represents the frequency of data points within
that range.
• Useful for visualizing density when many points overlap in a scatter
plot.
• A 3D histogram (or bivariate histogram) visually represents how
frequently pairs of attribute values (e.g., "weight" and "height") occur in a
dataset.
• Example: Histogram for Weight & Height
– X-axis: Represents weight.
– Y-axis: Represents height.
– Z-axis (bar height): Represents the frequency of occurrences for each (weight, height)
combination.
• This type of visualization is useful when analyzing the distribution of two
quantitative attributes.
• While 3D histograms help visualize frequencies of value combinations, they
have a downside—some bars can be hidden due to overlapping.
• A more effective way to analyze the relationship between two quantitative
attributes is using scatter plots.
What Is a Scatter Plot?
• A scatter plot is a 2D visualization where:
• X-axis represents one quantitative attribute (e.g., weight).
• Y-axis represents another quantitative attribute (e.g., height).
• Each dot represents an object in the dataset.
• It helps identify correlations, trends, and outliers between the two attributes.
Types of Correlations in Scatter Plots
• Scatter plots help us determine how two variables relate:
• Positive Correlation (⬆⬆): As one attribute increases, the other also increases.
• Negative Correlation (⬆⬇): As one attribute increases, the other decreases.
• No Correlation: No clear relationship between the attributes.
Covariance: Measuring the Relationship Between Two Attributes

• Covariance measures the direction of the relationship between two


quantitative attributes.
• It tells us whether they increase or decrease together and by how much.
Mathematical Formula for Covariance
• For two variables X and Y, the covariance is calculated as:
Interpretation of Covariance
Positive Covariance (Cov(X,Y)>0):
• As X increases, Y also increases (direct relationship).
• Example: Height and Weight (Taller people tend to
weigh more).
Negative Covariance (Cov(X,Y)<0):
• As X increases, Y decreases (inverse relationship).
• Example: Speed of a Car vs Travel Time (Higher speed
leads to lower travel time).
Zero Covariance (Cov(X,Y)≈0):
• No clear relationship between XXX and YYY.
• Example: Height and Favorite Color (No correlation).
Why Covariance Alone Is Not Enough?
• The magnitude of covariance depends on the
scale of the attributes, making it difficult to
interpret.
• That’s why we often normalize it using the
correlation coefficient.
Variance as a Special Case of Covariance
• Variance is just the covariance of a variable with
itself: Var(X)=Cov(X,X)
• It measures how much a single variable spreads
around its mean.
Pearson Correlation: A Normalized Measure of Relationship

• While covariance tells us whether two attributes move


together, its magnitude depends on the scale of the data.
• Pearson correlation, on the other hand, is a standardized
measure that is not affected by scale differences.
Interpreting Correlation with Scatter Plots
Three types of correlation between two variables, A and B:
• Positive Correlation (r>0)
– As A increases, B also increases.
– Example: Study time vs Exam scores → More study leads to
better scores.
– Points tend to form a rising straight line.
• Negative Correlation (r<0)
– As A increases, B decreases.
– Example: Speed vs Travel time → Faster speeds result in shorter
travel times.
– Points form a falling straight line.
• No Correlation (r≈0)
– No clear linear relationship between A and B.
– Example: Height vs Favorite Color → No meaningful connection.
– Points are scattered randomly.
Understanding Pearson Correlation Coefficient
• The Pearson correlation coefficient (r) quantifies the strength
and direction of a linear relationship between two variables.
Spearman's Rank Correlation

• Measures monotonic relationship (whether


one variable increases as the other does, not
necessarily at a constant rate)

rxiand ryi​ are the ranks of x and y.


rx​ˉ​ and ry​ˉ​ are the mean ranks.
srx​ and sry​​ are the standard deviations of ranks.
Basis for Assigning Ranks in Spearman's Rank Correlation
• In Spearman's Rank Correlation, ranks are assigned based on the
ascending order of the values.
• The smallest value gets rank 1, the second smallest gets rank 2, and so on.
Step-by-Step Process for Assigning Ranks:
• Sort the values in ascending order (smallest to largest).
• Assign ranks based on position.
• If two or more values are the same (ties), assign them the average rank of
the positions they would occupy.
Two Qualitative Attributes, at Least one of them Nominal
• When dealing with two qualitative attributes, where at least one is
nominal, we use contingency tables (also called cross-tabulation tables)
to analyze their relationship.
Structure of a Contingency Table
• A contingency table is a matrix of counts showing how often different
combinations of categorical variables occur.

• Rows: Represent categories of one attribute.


• Columns: Represent categories of the second attribute.
• Cells: Show the joint frequency (number of occurrences of that
combination).
• Row Totals: Sum of values across the row.
• Column Totals: Sum of values in each column.
• Grand Total: Total number of observations.
• This example demonstrates how contingency tables can be used to analyze
relationships between two qualitative attributes, in this case, "Gender" and
"Company Rating" (Good/Bad).
Mosaic Plots: A Visual Representation of Contingency Tables
• A mosaic plot is a graphical method for visualizing contingency tables, where
the size of each tile represents the relative frequency of each category
combination.
• It is useful for displaying relationships between two or more categorical
variables.

How Mosaic Plots Work


• Each rectangle's area represents the proportion of occurrences in that
category.
• The plot is divided vertically and horizontally according to the contingency
table’s row and column proportions.
• A larger rectangle means a higher frequency for that combination.
• The color may indicate statistical significance or another attribute.
Two Ordinal Attributes

• When working with two ordinal attributes, we can use several methods from
bivariate analysis, but with specific adjustments to account for the ordered nature of
the data.
Key Considerations for Two Ordinal Attributes
• Spearman’s Rank Correlation (𝜌)
– Used instead of Pearson correlation because it measures monotonic
relationships, not just linear ones.
– Works by ranking the values and calculating correlation based on their ranks.
• Scatter Plots with Jittering
– A standard scatter plot can be misleading because multiple values might
overlap, making it difficult to see the density of points.
– Jittering adds a small random variation to the points, making the distribution
more visible.
• Contingency Tables & Mosaic Plots
– Useful for summarizing relationships between ordinal variables.
– When used for ordinal data, categories must be in increasing order (e.g., "Poor"
→ "Average" → "Good" → "Excellent").
Descriptive Multivariate Analysis
• Multivariate analysis is essential when dealing with real-
world datasets, which often contain many attributes.

• In fields like biology, finance, and machine learning, analyzing


multiple variables simultaneously helps uncover patterns and
relationships that would not be visible in univariate or
bivariate analysis.

• As in univariate and bivariate analysis, frequency tables,


statistical measures and plots can be used or adapted for
multivariate analysis.
Multivariate Frequencies
• Multivariate frequency analysis extends univariate frequency analysis by considering
multiple attributes at once.
Types of Frequency Measures for Each Attribute
1. Absolute Frequency (Count)
• Number of times a specific value appears in the dataset.
• Example: How many people are exactly 170 cm tall?

[Link] Frequency (Proportion or Percentage)


• Computed by dividing absolute frequency by the total number of observations.
• Example: What percentage of people are 170 cm tall?

[Link] Cumulative Frequency (Running Total)


• Sum of absolute frequencies up to a given value.
• Example: How many people are 170 cm tall or shorter?

[Link] Cumulative Frequency (Cumulative Proportion)


• Cumulative frequency divided by total observations.
• Example: What proportion of people are 170 cm tall or shorter?
Multivariate Data Visualization
• When working with more than two attributes, traditional plots like histograms and
scatter plots become insufficient.

• However, several visualization techniques can extend bivariate plots or introduce


new approaches to represent complex data relationships.

• some of the previous plots can be extended to represent a small number of extra
attributes.

• new visualization approaches and techniques are continuously being created to


deal with new types of data, new approaches to results interpretation and new
data analysis tasks.

• Depending on the number of attributes, and the need to represent spatial and/or
temporal aspects of the data, different plots can be used.

• This section explores how multivariate data can be visually represented in different
ways and the main benefits of each of these alternatives.
• When the multivariate data has three attributes, or one can
only analyze three attributes from a multivariate data set, the
data can still be visualized in a bivariate plot, associating the
values of the third attribute with how each data object is
represented in the plot.

• If the third attribute is quantitative, the value can be


represented by the size of the object representation in the
plot.
• The plot is a bubble plot, which is an extension of a standard scatter plot
for three quantitative attributes.
Here's how the visualization works:
• X-axis (horizontal): Represents Weight
• Y-axis (vertical): Represents Height
• Bubble size: Represents Maximum Temperature
(Maxtemp)
• It allows us to visualize three attributes
simultaneously in a simple 2D space.
• The larger the bubble, the higher the value of the third attribute
(Maxtemp in this case).
• It provides an intuitive way to identify trends and relationships between
variables.
Representing a Qualitative Third Attribute in a Plot
• When the third attribute is qualitative (categorical) instead of
quantitative, we can represent it using:
✔ Color: Different colors for different categories
✔ Shape: Different marker shapes (e.g., circles, squares,
triangles)
• The number of colors or shapes will be the number of values
the attribute can assume.
• In classification tasks, color and shape are usually employed
to represent the class labels
• As an example, Figure shows two plots where the third attribute is
qualitative.
• On the right-hand size it represents each qualitative value as a different
shape.
• On the left-hand size, it represents each qualitative value by a different
color.
3D Plots for Multivariate Data
• A three-dimensional (3D) plot is useful when all three
attributes are quantitative because:
✔ Each axis (X, Y, Z) represents a numerical attribute.
✔It maintains the natural ordering of data values.
✔It helps visualize relationships between three variables in
space.
Multivariate Data Representation (More than 3 Attributes)
• When dealing with more than three attributes, we can extend visualization
techniques:
Methods to Represent Extra Attributes
✔ Size of Points → Represents a fourth attribute (useful for quantitative values).
✔ Color of Points → Encodes another attribute (categorical or numerical).
✔ Shape of Points → Differentiates categories (e.g., circles for males, squares for
females).
Challenges in Visualizing High-Dimensional
Data
• While 3D plots, color, and shape help represent
multiple attributes, they come with limitations:
Issues with High-Dimensional Visualizations
• Loss of Order & Magnitude → Color/shape works
for categories but doesn’t preserve numerical
relationships.

• Confusion & Overlapping → Too many attributes


in a single plot can make interpretation difficult.

• Scaling Issues → Some attributes (e.g., qualitative


values) don’t map naturally to size variations.
Parallel Coordinates Plot: A Powerful Tool for High-
Dimensional Data

• When dealing with more than four attributes,


traditional plots become cluttered.
• Parallel Coordinates Plot (PCP) is a great alternative
for visualizing multivariate quantitative data
efficiently.
• Each vertical axis represents a different attribute.
• Objects (data points) are represented as lines
connecting the attribute values across axes.
• Higher attribute values are placed higher on the
respective axis.
• Patterns, trends, and outliers become visible based
on line clustering and crossings.
Benefits of PCP
• Can handle many attributes in one
visualization.

• Preserves relationships between variables.

• Helps identify correlations and patterns in


high-dimensional data.
•It is a parallel coordinate plot for four of our contacts,
using three quantitative predictive attributes.

•The quantitative attributes occupy positions on the


vertical axes related to their values.

•Each object is represented by a sequence of lines


crossing the vertical axes at heights that represent the
attribute values.

•It is easy to see the values of the attributes for each


object.

•According to this plot, the attribute values of three of


the objects have a similar pattern, which is very
different from the profile of the fourth.

•The plot also shows the minimum and maximum


values for each attribute: the highest and the lowest
values on each vertical axis
• As we add more objects and more attributes, the lines tend to cross each
other, making the analysis more difficult.

• This figure also shows that qualitative attributes can be represented in


parallel coordinate plots.
• In this case, the qualitative attribute “gender” has been incorporated.
• Since it has has only two values, “M” and “F”, all the objects go to one of
two positions on the vertical axis
• Although the plot looks confusing, we can make the analysis of the data in
a parallel coordinate plot easier by assigning a color or style to each class
and using it for plotting the lines of corresponding objects.
• Thus, the sequences of lines for the objects from the same class will have
the same color or style.
• a modified version of the previous plot, using a solid line for contacts who
are good company and dotted lines for those who are bad company.

• Even with this modification, the analysis of the information in the plot is
still not easy.
• The ease of interpretation of these plots depends on the
sequence of the attributes used.

• If the lines from different objects keeping crossing each other,


it can be very difficult to extract information from the plot.

• Changing the order in which the attributes are plotted can


lead to fewer crossings.

• Each line in a parallel coordinate plot represents an object.

• If you do not have many objects and would like to look at each
of them individually, you can use another plot, known as the
star plot (it is also known as a spider plot or a radar chart).
•Figure 3.8 shows star plots for four quantitative
attributes (maxtemp, height, weight and years).

•To avoid the predominance of attributes with larger


values in the plot, all attributes have their values
normalized to the interval [0.0, 1.0].

•When the value of an attribute is close to 0.0, its


corresponding star will be too close to the center to be
seen.

•This is clear in the star labeled “Irene”. As can be seen in


Table 3.1, Irene has the lowest values for some of the
attributes.
• Qualitative attributes can also be represented in a star plot.
• However, since they have small numbers of values, points for qualitative
attributes will have few variations.
• We can also label each star in the star plot.
• Figure 3.9 shows two star plots for five attributes (maxtemp, height,
weight, years and gender), mixing quantitative and qualitative attributes.
• In the star plot on the left, each star is labeled with the contact’s name. In
the star plot on the right, each object is labeled with its class
• Taking advantage of the facility of human beings to recognize
faces, Herman Chernoff proposed the use of faces to
represent objects.
• an approach now referred to as Chernoff faces.
• Each attribute is associated with a different feature of a
human face.
• If the number of attributes is smaller than the number of
features, each attribute can be associated with a different
feature.
• Figure shows how our contacts data set can be represented by
Chernoff faces, using as attributes “maxtemp”, “height”,
“weight”, “years” and “gender”
Multivariate Extensions of Basic Statistical Measures
Location Multivariate Statistics
• To measure the location statistics when there are
several attributes , measure the location of each
attribute.

• Thus, the multivariate location statistical values can be


computed independently for each attribute.

• These values can be represented by a numeric vector


whose number of elements is equal to the number of
attributes.

• Since location statistics measure the central position of


data, we can compute them independently for each
attribute and represent them as a vector.
Dispersion Multivariate Statistics
• For multivariate statistics, dispersion statistics, such as
the amplitude, inter quartile range, mean absolute
deviation and standard deviation, can be independently
defined for each attribute
Statistical Methods for Evaluation
• Visualization is useful for data exploration and
presentation, but statistics is crucial because it
may exist throughout the entire Data Analytics
Lifecycle.
• Statistical techniques are used during the
– initial data exploration
– data preparation
– model building
– evaluation of the final models
– assessment of how the new models improve the
situation when deployed in the field.
Statistics Can Help Answer The Following Questions For Data
Analytics:

Model Building and Planning


• Key questions:
– What are the best input variables?
– Can the model predict the outcome?

Model Evaluation
• Evaluates:
– Model accuracy
– Performance compared to a guess
– Performance vs. other models

Model Deployment
• Considerations:
– Is the prediction sound?
– Does the model achieve the desired effect?
Hypothesis Testing
• Hypothesis testing is a fundamental statistical
method used to compare populations and assess
whether the difference between sample means is
statistically significant.
• Hypothesis method compares two opposite
statements about a population and uses sample data
to decide which one is more likely to be correct.
• To test this assumption, we first take a sample from
the population and analyze it and use the results of
the analysis to decide if the claim is valid or not.
Distributions of two samples of data
Steps of Hypothesis Testing

• Step 1: State your hypotheses


• Step 2: Collect and prepare data
• Step 3: Choose the appropriate statistical test
• Step 4: Calculate the test statistic and p-value
• Step 5: Make a decision
• Step 6: Present your findings
Step 1: State your hypotheses
• Null hypothesis (H0):
– The null hypothesis is the starting assumption in
statistics.
– This is the default assumption that there is no
effect or no difference.
– It's a statement of no change or no association.
– It says there is no relationship between groups.
– For Example A company claims its average
production is 50 units per day then here:
Null Hypothesis: H₀: The mean number of daily visits (μ) = 50.
• Alternative hypothesis (H1):
– The alternative hypothesis is the opposite of the null
hypothesis it suggests there is a difference between
groups.
– This is what you aim to prove.
– It's a statement indicating some effect or difference.
– like The company’s production is not equal to 50
units per day then the alternative hypothesis would
be:
H₁: The mean number of daily visits (μ) ≠ 50.
• Correctly stating the null hypothesis (H₀) and
alternative hypothesis (H₁) is crucial because
the entire hypothesis testing process depends
on them.
• If they are misstated, the conclusions drawn
from the test may be incorrect, leading to
invalid scientific or business decisions.
Example Null Hypotheses and Alternative Hypotheses
• Once a model is built over the training data, it
needs to be evaluated over the testing data to
see if the proposed model predicts better than
the existing model currently being used.

• The null hypothesis is that the proposed model


does not predict better than the existing model.

• The alternative hypothesis is that the proposed


model indeed predicts better than the existing
model.
• When evaluating a model, sometimes it needs to be
determined if a given input variable improves the model.

• In regression analysis , for example, this is the same as


asking if the regression coefficient for a variable is zero.

• The null hypothesis is that the coefficient is zero, which


means the variable does not have an impact on the
outcome.

• The alternative hypothesis is that the coefficient is


nonzero, which means the variable does have an impact
on the outcome.

• A common hypothesis test is to compare the means of


two populations
Difference of Means
• Hypothesis testing is a statistical method used
to determine whether two populations,
denoted as P1 and P2, have significantly
different means.
• The approach is based on comparing sample
means, and , to infer about the
population means μ1 and μ2.
• To formally test the difference, we use Student’s t-test or Welch’s t-test, depending on
whether the variances of the populations are equal.
Keywords
• Level of significance:
– It refers to the degree of significance in which we
accept or reject the null hypothesis.
– 100% accuracy is not possible for accepting a
hypothesis so we select a level of significance.
– This is normally denoted with α and generally it is
0.05 or 5% which means your output should be
95% confident to give a similar kind of result in
each sample.
• Critical value:
– Critical value is a boundary or threshold that helps
you decide if your test statistic is enough to reject
the null hypothesis
• Degrees of Freedom:
– It is defined as the maximum number of
independent values that can vary in a sample
space.
– The degree of freedom is generally calculated
when we subtract one from the given sample of
data.
– Degrees of freedom are very helpful for ensuring
the validity of chi-square tests, t-tests, high f-tests,
and others.
Example 1: Choosing Meals
• Degree of Freedom can easily be understood with the help of the
following example. Suppose you have three packets of A, B, and C
of food to eat in a day.

• For Breakfast: You can have any of the three packets (A, B, and C),
but you choose to eat packet A.

• For Lunch: As you have only two packets left (B, C) you have two
choices and you choose B.

• For Dinner: You have only one packet left i.e. C, so you are forced to
eat that one only.

• So, we have two levels of freedom to choose our food for breakfast
in a day, so the degree of freedom in this case is two, but for lunch,
we have one freedom so the degree of freedom in this case is one
and for dinner, we have no choices so the degree of freedom is one.
Student’s t-test
• The t-test is a statistical test used to compare
the means of two groups.
• It helps determine if differences between the
means are statistically significant.
• Assumptions:
– Data is normally distributed.
– Variances are equal (for independent t-test with
equal variance).
– Samples are independent or paired, depending on
the type of test.
• Suppose n1 and n2 samples are randomly and
independently selected from two populations,
pop1 and pop2 , respectively.
• If each population is normally distributed with
the same mean 𝜇1 = 𝜇2 and with the same
variance, then T (the t-statistic), follows a t-
distribution with n1+n2-2 degrees of freedom
(df).
• The t-distribution is similar in shape to the
normal distribution but has heavier tails, which
account for more variability in small sample sizes.

• As the degrees of freedom (df) increase


(approaching 30 or more), the t-distribution
becomes almost identical to the normal
distribution.
Welch’s t-test

• Welch’s t-test is a statistical test used to


compare the means of two independent
groups.

• It is an adaptation of the Student’s t-test but


does not assume equal variances between
groups.
When to Use Welch’s t-Test?
• When comparing the means of two independent groups.

• When the assumption of equal variances is violated.

• When sample sizes are unequal.

Key Assumptions

• The two samples are independent.

• The data are approximately normally distributed.

• The variances of the two groups may be unequal.


Welch’s t-Test Formula

Degrees of Freedom Calculation


Steps to Perform Welch’s t-Test

• State the null and alternative hypotheses.


• Calculate the test statistic (t-value).
• Determine the degrees of freedom.
• Find the critical value or p-value.
• Compare with the significance level () and
interpret the results.
Advantages of Welch’s t-Test
• More reliable when variances are unequal.
• Does not require equal sample sizes.
• More robust than Student’s t-test.
Limitations
• Assumes normality of data.
• Can be less powerful than Student’s t-test if
variances are equal.
• Interpretation remains similar to other t-tests.
Wilcoxon Rank Sum Test: A Non-Parametric
Alternative

• The Wilcoxon Rank Sum Test is a non-


parametric test used to compare two
independent groups.

• It is an alternative to the independent t-test


when the assumption of normality is violated.

• Also known as the Mann-Whitney U test.


When to Use the Wilcoxon Rank Sum Test?
• When comparing the central tendencies of two
independent groups.
• When data is not normally distributed.
• When data is ordinal or has outliers.
Wilcoxon Rank Sum Test Procedure
• Combine the two samples and rank all values.
• Assign ranks to each observation.
• Compute the rank sum for each group.
• Calculate the test statistic (W).
• Compare W with the critical value or compute a
p-value.
Wilcoxon Rank Sum Test Statistic
• Example:
• Two independent groups:
– Group A: [5, 8, 12]
– Group B: [7, 10, 15]
• Rank all values: [1, 2, 3, 4, 5, 6].
• Sum ranks for each group.
• Compare W with critical values
Step 1: Arrange the Data
• We have two independent groups:
• Group A: [5, 8, 12]
• Group B: [7, 10, 15]
Step 2: Rank the Combined Data
Type I Error
• A type I error appears when the null hypothesis (H0) of an experiment is true, but still, it
is rejected.

• It is stating something which is not present or a false hit.

• A type I error is often called a false positive (an event that shows that a given condition
is present when it is absent).

• In words of community tales, a person may see the bear when there is none (raising a
false alarm)
– where the null hypothesis (H0) contains the statement: “There is no bear”.

• The type I error significance level or rate level is the probability of refusing the null
hypothesis given that it is true.

• It is represented by Greek letter α (alpha) and is also known as alpha level.

• Usually, the significance level or the probability of type i error is set to 0.05 (5%),
assuming that it is satisfactory to have a 5% probability of inaccurately rejecting the null
hypothesis.
Type II Error
• A type II error appears when the null hypothesis is false but mistakenly fails to be refused.

• It is losing to state what is present and a miss.


• A type II error is also known as false negative (where a real hit was rejected by the test and is
observed as a miss), in an experiment checking for a condition with a final outcome of true or
false.

• A type II error is assigned when a true alternative hypothesis is not acknowledged.


• In other words, an examiner may miss discovering the bear when in fact a bear is present
(hence fails in raising the alarm).

• Again, H0, the null hypothesis, consists of the statement that, “There is no bear”, wherein, if a
wolf is indeed present, is a type II error on the part of the investigator.

• Here, the bear either exists or does not exist within given circumstances, the question arises
here is if it is correctly identified or not, either missing detecting it when it is present, or
identifying it when it is not present.

• The rate level of the type II error is represented by the Greek letter β (beta) and linked to the
power of a test (which equals 1−β).
Medical Testing:
Suppose a medical test is designed to diagnose a particular disease.

The null hypothesis (H0) is that the person does not have the disease, and the alternative hypothesis
(H1) is that the person does have the disease.

A Type I error occurs if the test incorrectly indicates that a person has the disease (rejects the null
hypothesis) when they do not actually have it.

A Type II error occurs if the test incorrectly indicates that a person does not have the disease (fails to
reject the null hypothesis) when they actually do have it

You might also like