0% found this document useful (0 votes)
6 views50 pages

Quantitative vs. Qualitative Data Explained

The document outlines the differences between quantitative and qualitative data, emphasizing that quantitative data is measurable and numerical while qualitative data is descriptive and observational. It also discusses primary and secondary data, highlighting that primary data is original and specific to the researcher's needs, whereas secondary data is pre-existing and less controlled. Furthermore, it covers descriptive and inferential statistics, detailing their roles in data analysis, including measures of central tendency, variability, and various statistical tools for hypothesis testing and confidence intervals.

Uploaded by

yashkamra
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views50 pages

Quantitative vs. Qualitative Data Explained

The document outlines the differences between quantitative and qualitative data, emphasizing that quantitative data is measurable and numerical while qualitative data is descriptive and observational. It also discusses primary and secondary data, highlighting that primary data is original and specific to the researcher's needs, whereas secondary data is pre-existing and less controlled. Furthermore, it covers descriptive and inferential statistics, detailing their roles in data analysis, including measures of central tendency, variability, and various statistical tools for hypothesis testing and confidence intervals.

Uploaded by

yashkamra
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Quantitative vs.

Qualitative data
Quantitative Qualitative

 Things that are measurable and can be  Qualitative data is usually not easily
expressed in numbers or figures, or using measurable as quantitative and can be gained
other values that express quantity through observation or open-ended survey or
interview questions.

 Quantitative data is usually expressed in  As quantitative data collection methods usually


numerical form and can represent size, length, are rather concerned with words, sounds,
duration, amount, price, and so on thoughts, feelings, and other non-quantifiable
data, it allows a greater depth of
understanding.

 Quantitative data is most likely to provide  Qualitative research is most likely to provide
answers to questions such as who? when? answers to questions such as “why?” and
where? what? and how many? “how?”

 Quantitative survey questions are in most  Qualitative data collection methods are most
cases closed-ended, thus making the likely to consist of open-ended questions
answers easily transformable into and descriptive answers and little or no
numbers, charts, graphs, and tables numerical value
Primary and Secondary Data
 The term primary data refers to the data originated by the researcher
for the first time. Secondary data is the already existing data, collected by
the investigator agencies and organizations earlier.
 Primary data is a real-time data whereas secondary data is one
which relates to the past.
 Primary data is collected for addressing the problem at hand while
secondary data is collected for purposes other than the problem at hand.
 Primary data collection is a very involved process. On the other
hand, secondary data collection process is rapid and easy.
 Primary data collection sources include surveys,
observations, experiments, questionnaire, personal interview, etc. On
the contrary, secondary data collection sources are government
publications, websites, books, journal articles, internal records etc.
Contd

 Primary data collection requires a large amount of resources
like time, cost and manpower. Conversely, secondary data is
relatively inexpensive and quickly available.
 Primary data is always specific to the researcher’s needs, and
he controls the quality of research. In contrast, secondary data is
neither specific to the researcher’s need, nor he has control over
the data quality.
 Primary data is available in the raw form whereas secondary data
is the refined form of primary data. It can also be said that
secondary data is obtained when statistical methods are applied to
the primary data.
 Data collected through primary sources are more reliable
accurate
and as compared to the secondary sources.
Statistics: Descriptive and
Inferential Statistics plays a main role in the field of
research. It helps us in the collection, analysis
and presentation of data.
Statistics is a branch of mathematics dealing
with the collection, analysis, interpretation,
and presentation of masses of numerical data.
 It is basically a collection of quantitative data.
Sample population which is selected randomly
for the study. The sample should be selected
such that it represents all the characteristics of
the population.
 Any group of data, which includes all the data you are interested in, is
called a population. A population can be small or large, as long as it
includes all the data you are interested in.
Statistics: Descriptive and
Inferential
Types of Statistics –
 Theoretical Statistics
 Applied Statistics
Descriptive Statistics
 It is a term given to the analysis
of data that helps to describe, show,
and summarize data in a meaningful
way.
It is a discipline that
describes the important characteristics
quantitatively
of the dataset.
 It is a simple way to describe our data.
 Descriptive statistics is very important
to present our raw data
ineffective/meaningful way using
numerical calculations or graphs or
tables.
 This type of statistics is applied
on already known data.
Descriptive
Statistics
 For example, if we had the results of 100 pieces of
students' coursework, we may be interested in the overall
performance of those students. We would also be interested in
the distribution or spread of the marks that can be done using
descriptive statistics.
There are two general types of statistic that are used to describe
data:
 o Measure of Central Tendency
 o Measure of Variability/Spread
Descriptive Statistics
 Central tendency: Use the mean or the median to locate the center of the
dataset. This measure tells you where most values fall.
 Dispersion: How far out from the center do the data extend? You can use
the range or standard deviation to measure the dispersion. A low
dispersion indicates that the values cluster more tightly around the
center. Higher dispersion signifies that data points fall further away from the
center. We can also graph the frequency distribution.
 Skewness: The measure tells you whether the distribution of values is
symmetric or skewed.
Inferential
Statistics  Inferential statistics takes data from a sample and
makes inferences about the larger population from
which the sample was drawn.
 Goal of inferential statistics is to draw conclusions
from a sample and generalize them to a population,
we need to have confidence that our
sample
accurately reflects the population.
 Random sampling allows us to have confidence that
the sample represents the population.
 Consequently, when you estimate the properties of
a population from a sample, the sample statistics
are unlikely to equal the actual population
value exactly.
 The difference between the sample statistic and the
population value is the sampling error.
Inferential Statistics
To Obtain a Representative Sample
1. Make sure you use a random sampling method.
 a. A simple random sample
 b.A systematic random sample
 c. A cluster random sample
 d.A stratified random sample

2. Make sure your sample size is large enough.


Inferential Statistics
TOOLS/Measures
• Hypothesis tests use sample data answer questions like the following:
1) Is the population mean greater than or less than a particular value?
2) Are the means of two or more populations different from each other?
Example: effectiveness of a new medication. After all, we don’t want to use
the medication if it is effective only in our specific sample. Instead, we need
evidence that it’ll be useful in the entire population of patients. Hypothesis
tests allow us to draw these types of conclusions about entire populations.

• Confidence intervals incorporate the uncertainty and sample error to create a


range of values in which actual population value is like to fall within.
Example: we draw a random sample from some population and calculate the
mean height as 181 cm. Now, a confidence interval of [176 186] indicates that
we can be confident that the real population mean falls within this range.
Inferential Statistics
TOOLS/Measures
• Regression Analysis: describes the relationship between a set of independent
variables(height) and a dependent variable(weight).
We have sufficient evidence to conclude that this relationship exists in the
population rather than just our sample.
Some differences to remember!
DESCRIPTIVE STATISTICS:
MEASURES OF CENTRAL
TENDENCY
MEAN, MEDIAN, MODE FOR UN-GROUPED data

mean = (sum of all the observations/total number of observations)


DESCRIPTIVE STATISTICS:
MEASURES OF CENTRAL
TENDENCY
MEAN, MEDIAN, MODE FOR UN-GROUPED data
Median

Definition: The median is the middle value in a sorted list of numbers, dividing the dataset into two equal halves.
Importance: It's a measure of central tendency that is less affected by outliers than the mean.
2. Median Calculation for Odd Number of Data Points:
Scenario: When the dataset has an odd number of observations.

Median=Value at position ((𝑛+1)/2)


Formula:

Median Calculation for Even Number of Data Points:


Scenario: When the dataset has an even number of observations.
Formula: Median=(Value at position (n/2)+Value at position (n/2+1)/2)2
Example: (Show a sorted dataset with an even number of values and calculate the median)
DESCRIPTIVE STATISTICS:
MEASURES OF CENTRAL
TENDENCY
MEAN, MEDIAN, MODE FOR UN-GROUPED data

MEAN – 61.3
MEDIAN – 61
MODE - 62
MEAN FOR GROUPED
DATA
Seconds Frequency

51-55

56-60

61-65

66-70
MEAN FOR GROUPED
DATA
Yes, you can use
L=60.5L = 60.5L=60.5
as the lower boundary
of the median class if
you interpret the class
intervals as continuous
rather than discrete.

MEAN – 61.3
MEDIAN – 61
MODE - 62
But the actual Mode may not even be
in that group! Or there may be more
than one mode. Without the raw data
we don't really know.
MEAN – 61.3
MEDIAN – 61
MODE - 62
TIME TO PRACTISE…
AGE EXAMPLE
• Age is a special case.
• When we say "Sarah is 17" she stays
"17" up until her eighteenth
birthday.
She might be 17 years and 364 days
old and still be called "17".
• This changes the midpoints and
class boundaries
• Example: The ages of the 112
people who live on a tropical island
are grouped as follows ->

Note: Boundary value will be starting


of interval value only
DESCRIPTIVE STATISTICS:
MEASURES OF
• SPREAD/VARIABILITY
The measure of variability is the statistical summary,
which represents the
dispersion within the datasets. On the other hand, the measure of central
tendency defines the standard value.

• Statisticians use measures of variability to check how far the data points
are going to fall from the given central value.

• The lower dispersion value shows the data points will be grouped nearer to
the center. The higher dispersion value shows the data points will be
clustered further away from the center.

• Four measures of Variability:


1. Range
2. Quartile
3. Standard Deviation
4. Variance
DESCRIPTIVE STATISTICS: RANGE
• It is used to know about the spread of the data from the least to
the most value within the distribution.

• Subtract the least value from the greatest value of the given dataset.

Suppose you have 5 data points as:

Data (minutes) 10 25 5 35 40
It is clear that 40 is the highest value and 5 is the lowest value. Therefore,

=> R = H-L => 40-5 => 35

The range of the data is 35 minutes.


DESCRIPTIVE STATISTICS: RANGE
• dataset 1 has a range of 20 – 38 = 18
while
dataset 2 has a range of 11 – 52 = 41.

• Dataset 2 has a broader range and,


hence, more variability than dataset 1

Note: the outliers can influence the range.


Moreover, the range does not give information
about value distribution.
DESCRIPTIVE STATISTICS:
INTERQUARTILE RANGE
• The IQR (interquartile range) provides the middle spread of the distribution.
• It is calculated by third quartile minus first quartile.
• We can divide the data into quarters. Statisticians refer to these quarters as
quartiles and denote them from low to high as Q1, Q2, and Q3.
• Interquartile range = Upper Quartile – Lower Quartile = Q3 – Q1
DESCRIPTIVE STATISTICS:
INTERQUARTILE RANGE

The range is 39 – 20
= 19

Note: the interquartile range is excellent for skewed


distributions,
DESCRIPTIVE STATISTICS:
STANDARD DEVIATION
• The SD is the mean of variability that tells how far the score is from the
average.
• It means the more the SD, the more variable data set would be.
• The standard deviation is just the square root of the variance.
DESCRIPTIVE STATISTICS:
STANDARD DEVIATION
Suppose you have 5 data points, and you have to calculate SD.

Deviation Squared Divide the Standard


Data from average Deviation addition Deviation

s = √1350 =
36.74
As we are The standard
70 -70 = 0 dealing with deviation of
70 110 – 70 = 40 1600 the sample,
110 400 the given data
50 – 70 = (- we need to is 36.74.
50 20) 2500 use n – 1.
20 900 It implies that
20 – 70 = (- n–1=5–1 score
100 50) Average of => 4
Average = 70 the square = deviation
100 – 70 = 30 5400/4 => away from
5400 1350 the 36.74
points.
DESCRIPTIVE STATISTICS:
VARIANCE
• Variance is the standard deviation’s square.
• The variance shows the degree of spread within the data sets. The larger
the variance, the larger the data spread.
DESCRIPTIVE STATISTICS:
VARIANCE
• Variance is the standard
deviation’s
square.
• The variance shows the degree of
spread within the data sets.
The larger the variance, the
larger the data spread.
EXAMPLE
DESCRIPTIVE STATISTICS:
VARIANCE
• The measure of asymmetry in a probability distribution
Skewness. It can either be positive, negative or undefined.
is defined by

• Positive Skew — This is the case when the tail on the right side of the curve is
bigger than that on the left side. For these distributions, mean is greater
than the mode.
• Negative Skew — This is the case when the tail on the left side of the curve is
bigger than that on the right side. For these distributions, mean is
smaller than the mode.
DESCRIPTIVE STATISTICS:
VARIANCE
The most commonly used method of calculating Skewness is:

If the skewness is zero, the distribution is symmetrical. If it is negative, the


distribution is Negatively Skewed and if it is positive, it is Positively
Skewed.
DESCRIPTIVE STATISTICS:
KURTOSIS
Kurtosis describes the whether the data is light tailed (lack of outliers) or heavy
tailed (outliers present) when compared to a Normal distribution. There are three
kinds of Kurtosis:
• Mesokurtic — This is the case when the kurtosis is zero, similar to the
normal distributions.
•Leptokurtic — This is when the tail of the distribution is heavy (outlier
present) and
kurtosis is higher than that of the normal distribution.
• Platykurtic — This is when the tail of the distribution is light( no outlier) and
kurtosis is lesser than that of the normal distribution.
Distribution: Different types of
Visualizations
1. Histograms
2. Bar Chart
3. Pie Chart
4. Scatter Graph
5. Line Graph
6. Box and Whisker Plot
Visualization
Histogr
am
[Link]
mislam/statistics-101-descriptive-
and-inferential#desc

[Link]
datasets/rashikrahmanpritom/
heart-attack-analysis-prediction-
dataset
import numpy as np
import pandas as pd
from scipy import stats
import statistics

You might also like