Statistics
Statistics are everywhere
Introduction of statistics
Descriptive and inferential statistics
Population and sample
Sampling methods
1. What is Statistics? Statistics is the studies of how to collect, organize, analyze, and interpret numerical information from data.
Descriptive statistics involves methods of organizing, picturing and summarizing information from data. Inferential statistics involves
methods of using information from a sample to draw conclusions about the population. Some consider statistics to be a
mathematical body of science pertaining to the collection, analysis, interpretation or explanation, and presentation of data, while
others consider it a branch of mathematics concerned with collecting and interpreting data. Because of its empirical roots and its
focus on applications, statistics is usually considered to be a distinct mathematical science rather than a branch of mathematics.
and Basic Concepts of Statistics
Statistics is the science of conducting studies to collect, organize, summarize, analyze, and draw conclusions from data.
Uses of Statistics in Different Fields:
Agriculture, Biology, Business, Demography Economics, Education, Engineering, . Entertainment, Health, Medicine, Research
Qualitative & Quantitative Variables Qualitative Variables describe a certain type of information without using numbers.
Examples: Color, Taste, Occupation, Gender, Preference to eating vegetables.
Qualitative & Quantitative Variables
Quantitative Variables measure or identify an information using numeric scales.
Examples: . Height . Number of sibling(s), Speed of a car, Temperature, Number of students in a classroom,Amount of shirt in the drawer.
Quantitative Variables can be classified as:
Discrete variable whose values can be counted using integral values Continuous variable can assume any numerical value over an interval or
intervals.
Descriptive vs Inferential Statistics
Population is the entirety of the group including all the members that forms a set of data Sample contains a few members of the population.
Samples were taken to represent the characteristics or traits of the population.
Descriptive Statistics used to say something or describe a set of information collected. It can also be represented with graphs.
Common tools are: Measures of Central Tendency • Measure of Variability
Ex: In a Math test, 32 out of 40 students were able to receive a passing mark. The average score of the class is 82 out of 100.
Inferential Statistics
Inferential Statistics used to say something about a larger group (population) using information collected from a small part of that population
(sample). Common tools are:
• Hypothesis Testing • Regression Analysis
Ex: In a sample survey conducted, 65% of Filipino Generation Z prefer to drink milk tea than coffee while only 34% of Filipino Millennials prefer
to drink milk tea than coffee.
Measurement Scales
Measurement is the process of assigning value to a variable.
There are four levels of scales of measurement:
Ratio, Interval,Ordinal, Nominal,Numerical,Categorical
Nominal Level
Observations can be named without particular order or ranking imposed on the data. Words, letters, and even numbers are used to classify the
data.
Example: Gender M – Male , F -Female -Types of Electric Consumption
/1- Residential ,2-Commercial,3 – Industrial,/4-Government
2) Ordinal Level
Describes ranking or order. The difference or ratio between two rankings may not always be the same.
Example:
Competition Placement
1-Champion 2-1st Runner-up 3-2nd Runner-up 4-3rd Runner-up
Level of Satisfaction
1-Very Satisfiede 2-Satisfied 3-Unsatisfied 4- Very Unsatisfied
3) Interval Level
Indicates an actual amount (numerical). The order and the difference between the variables can be known. Its limitation is it has no "true zero".
Example: Temperature:60 °C, 20 °C, 10 °C,-15 °Ca
4) Ratio Level
It has the same properties as the interval level. The order and difference can be described. Additionally, it has a true zero and the ratio
between two points has a meaning
Example:Mass 80 kg,40 kg,10 kg,O kg
Categorical
- Male/female
- Qualitative data(characteristics )
Numerical
-age, height,weight, any measure that use number
Sample/sampling methods
-you can draw a conclusion ,selecting or process out of large group.
You can generalized soemthing (e.g brgy,majority are farmers)
Output(samples) process(sampling)
Sampling (valid alternative to census) when doing sampling
Survey entire population in impractical, questions; how they represent the population
-budget constraints restrict data collection how similar and significant
Time constraints restrict data collection
Population total quantity things(cases) sampling frame
It’s an object , organizations, people or events you certainly choose a certain group.
Sampling methods
Probability sampling and non probability sampling
Probability- reliable representation of the whole population.
Non-probability – refers to relying on the judgement od the researcher.
Probabilty – equal chance, use statistical theory(sloving formula), predict that much to the population.
Four probability sampling
Simple random sampling cluster random sampling
Assigning numbers to an individual’s(samples) analyzed a particular population
E.g JL deguzman 8. E.g district 1-9 cluster (pop. Of non-gov. Employee)
Stratified random sampling Systematic sampling
Large population can devided into small group. There’s pattern e.g 1-10, then pick number 6, so you count per 6 head then choose it
Method- classify,set. ( Then as sample (needed a list and number also)
Types of Non- probability
Convenience sampling- ease of access/obtain but not provide reliable.
Judgemental sampling- you sampling because you’re searching for something.,( has a purpose)
(e.g) PSU BC- BPA- WORKING STUDENT ETC.
Snowball sampling- referring they use population of people e who hard to find (e.g)( drug user,)
Quota sampling – may hinahangad na qouta.
SAMPLING MEANS- ALIGN TO SUBJECT TO THE RESEARCH
Type of variablesDependent and Independent VariablesAn independent variable, sometimes called an experimental or predictor variable, is a
variable that is being manipulated in an experiment in order to observe the effect on a dependent variable, sometimes called an outcome
[Link] dependent variable is simply that, a variable that is dependent on an independent variable(s). For example, in our case the test
mark that a student achieves is dependent on revision time and intelligence.
Sampling techniquesIn population data, the variable is from every individual of interest. In sample data the variable is only from some of the
individuals of interest:Random SampleA simple random sample of n measurements from a population is one selected in such a manner that:
Every sample of size n from the population has equal probability of being selected, and Every member of the population has equal probability
of being included in the samples. Stratified SamplingStratify divide the population by a common characteristic such as age, class, gender and
take a random sample of each stratum [often in accordance to their per cent of the population].for examples, divide the population by race and
survey a number of randomly selected individuals from each race often in the same proportion they occur in the overall population.
Systematic SamplingUsed when the elements of the population are arranged in a natural sequential order Cluster SamplingUsed extensively by
government and research organizations. Randomly select a sample of preexisting sections or clusters often geographic sections. Every member
of the cluster is included in the sample or survey. For examples, randomly select 30 schools and survey every student in each school.
Convenience SamplingUsing data that are conveniently and readily obtained. For examples, walk outside and survey the first 100 people that
will talk.
5. Experimental and Observational Studies
Experimental studiesThe experimental group is actually given the treatment and the control group is not given the treatment. In an experiment
investigators apply treatments to experimental units of people, animals, and plots of land and then proceed to observe the effect of the
treatments on the experimental units. Ina randomized experiment investigators control the assignment of treatments to experimental units
using a chance mechanism (like the flip of a coin or a computer's random number generator
Observational Studies In an observational study, observations and measurements of individuals are conducted in a way that does not change the
response or the variable being measured. In an observational study investigators observe subjects and measure variables of interest without
assigning treatments to the subjects. The treatment that each subject receives is determined beyond the control of the investigator.
Chapter 3
Population in Research
• It does not necessarily mean a number of people, it is a collective term used to describe the total quantity of things (or cases) of the type which
are the subject of your study.
• So a population can consist of certain types of
objects, organizations, people or even events.
The need to sample
Sampling- a valid alternative to a census when;
• A survey of the entire population is impracticable Budget constraints restrict data collection Time constraints restrict data collection
• Results from data collection are needed quickly.
a survey, the question inevitably arises:
• how representative is the sample of the whole
population, in other words;
• how similar are characteristics of the small group of cases that are chosen for the survey to those of all of the cases in the whole group?
Sampling Frame when this population, there will probably be only certain groups that will be of interest to your study, this selected category is
your sampling frame. Relationship between Population, Sampling Frame and Sampling frame
Probability Sampling
• It is a sampling technique in which sample from a larger population are chosen using a method based on the theory of
probability.
• For a participant to be considered as a probability sample, he/she must be selected using a random selection.
• The most important requirement of probability sampling is that everyone in your population has a known and an equal chance of getting
selected.
Probability sampling uses statistical theory to select randomly, a small group of people (sample) from an existing large population and then
predict that all their responses together will match the overall population.
Types of Probability Sampling
Four main techniques used for a probability sample:
> Simple random
> Stratified random
>Cluster
>Systematic
Simple random sampling
suggests is a completely random method of selecting the sample. This sampling method is as easy as assigning numbers to the individuals
(sample) and then randomly choosing from those numbers through an automated process.
Stratified Random sampling
• involves a method where a larger population can be divided into smaller groups, that usually don't overlap but represent the entire
population together. While sampling these groups can be organized and then draw a sample from each group separately.
(A common method is to arrange or classify by sex, age, ethnicity and similar ways.)
Cluster random sampling
way to randomly select participants when they are geographically spread out. Cluster sampling usually analyzes a particular population in
which the sample consists of more than a few elements,
(for example, city, family, university etc.)
The clusters are then selected by dividing the greater population into various smaller sections.
Stratified & Cluster Sampling
Stratified
•Population divided into few subgroups
• Each subgroup has many elements in it.
• Subgroups are selected according to
some criterion that is related to the
variables under study.
Cluster
Homogeneity within subgroups
• Heterogeneity between subgroups
Choice of elements from within each
subgroup Cluster.
• Population divided into many subgroups
• Each subgroup few elements in it. Subgroups are selected according to some criterion of ease or availability in data collection.
• Heterogeneity within subgroups
• Homogeneity between subgroups
• Random choice of subgroups
Convenience Sampling
a non-probability sampling technique used to create sample as per ease of access, readiness to be a part of the sample, availability at a given
time slot or any other practical specifications of a particular element. Convenience sampling involves selecting haphazardly those cases that
are easiest to obtain for your sample, such as the person interviewed at random in a shopping center for a television program.
Judgmental Sampling
In the judgemental sampling, also called purposive sampling,
Line sample members are chosen only on the basis of the researcher's knowledge and judgment.
• It enables you to select cases that will best enable you to answer your research question(s) and to meet your objectives.
Snowball Sampling
Snowball sampling method is purely based on referrals and that is how a researcher is able to generate a sample. Therefore this method is
also called the chain-referral sampling method.
• This sampling technique can go on and on, just like a snowball increasing in size (in this case the sample size) till the time a researcher has
enough data to analyze, to draw conclusive results that can help an organization make informed decisions.
Quota Sampling
Selection of members in this sampling technique happens on basis of a pre-set standard. In this case, as a sample is formed on basis of specific
attributes, the created sample will have the same attributes that are found in the total population. It is an extremely quick method of
collecting samples.
• Quota sampling is therefore a type of stratified sample in which selection of cases within strata is entirely non-random Quota Sampling.
Probability Sampling
As sampling technique in which sample from a larger Population are chosen using a method based on the theory of probability
For a participant to be considered as a probability sample, he/she must be selected using a random selection.
•The most important requirement of probability sampling that everyone in your population has a known and an equal chance of getting
selected.
Probability sampling uses statistical theory to select randomly, (sample) from an existing large a small group of people (sample) population and
then predict that all their responses together will match the overall population.
Homogenous- all cases are similar ([Link] of beer on production line)
Stratified- contain strata or layers (ex. People with different levels of income ,low medium, high )
Proportional stratified - contains strata of known proportion (ex. Percentages of different nationalities of students in a university
Group by type -contains distinctive group (ex. Of apartment building, towers, slabs,villas, tenement blocks)
Group by location - different groups according to where they are (ex. Animals in different habitats , desserts, equatorial forest savannah, tundra)
Measures of Central tendency, despersion and position
MEAN, MEDIAN, MODE
In statistics, mean, median, and mode are measures of central tendency, meaning they represent the typical or central value of a data set.
1. Mean: The mean, often referred to as the average, is calculated by adding up all the values in a data set and then dividing by the number of
values. It's the sum of all the values divided by the total number of values.
2. Median: The median is the middle value in a sorted list of numbers. If there is an odd number of values, the median is simply the middle
number. If there is an even number of values, the median is the average of the two middle numbers.
3. Mode: The mode is the value that appears most frequently in a data set. It's possible for a data set to have one mode, more than one mode
(multimodal), or no mode at all (when all values occur with the same frequency).
These three concepts help statisticians understand the central tendency and distribution of data, providing insights into the characteristics of
a dataset.
Range and STANDARD DEVIATION
In statistics, range and standard deviation are measures of variability or dispersion within a dataset.
1. Range: The range is the difference between the highest and lowest values in a dataset. It gives a sense of how spread out the values are. To
calculate the range, you subtract the lowest value from the highest value.
2. Standard Deviation: The standard deviation measures the average distance of each data point from the mean of the dataset. It indicates the
extent to which data points deviate from the mean. A low standard deviation suggests that the data points tend to be close to the mean, while a
high standard deviation indicates that the data points are spread out over a wider range. It's calculated by taking the square root of the average
of the squared differences between each data point and the mean.
In statistics, the coefficient of variation (CV) is a measure of relative variability, which is used to compare the spread of data sets with different
means and units of measurement. It is calculated by dividing the standard deviation of the data set by the mean, and then multiplying by 100 to
express the result as a percentage.
The formula for the coefficient of variation (CV) is:
The coefficient of variation is particularly useful when comparing the variability of data sets with different scales or units, as it provides a
standardized measure of dispersion relative to the mean. A lower CV indicates less variability relative to the mean, while a higher CV suggests
greater variability relative to the mean.
In percentiles, deciles, and quartiles
(are measures used to divide a dataset into equal parts, providing insight into the distribution of values).
1. Percentiles: Percentiles divide a dataset into 100 equal parts. For example, the 25th percentile (also known as the first quartile) represents the
value below which 25% of the data falls. Similarly, the 50th percentile (median) represents the value below which 50% of the data falls, and so
on.
2. Quartiles: Quartiles divide a dataset into four equal parts. The first quartile (Q1) is the same as the 25th percentile, the second quartile (Q2) is
the median (50th percentile), and the third quartile (Q3) is the 75th percentile. Quartiles provide insight into the spread and skewness of the
data.
3. Deciles: Deciles divide a dataset into ten equal parts. The first decile represents the 10th percentile, the second decile represents the 20th
percentile, and so on up to the ninth decile, which represents the 90th percentile. Deciles provide finer granularity in understanding the
distribution of data compared to quartiles.
In statistics, the standard normal distribution, also known as the Gaussian distribution or bell curve, is a specific type of probability
distribution. It is characterized by a symmetric, bell-shaped curve where the mean, median, and mode are all equal and located at the center of
the distribution.
The standard normal distribution has a mean of 0 and a standard deviation of 1. This means that the curve is centered at 0 on the x-axis, and
the spread of the curve is such that about 68% of the data falls within one standard deviation of the mean, about 95% falls within two standard
deviations, and about 99.7% falls within three standard deviations.
The equation for the standard normal distribution is given by the probability density function:
The standard normal distribution is widely used in statistical analysis and hypothesis testing, as many statistical methods assume that the data
follows a normal distribution.
MESOKURTIC (KURTOSIS= 3)
Same as the normal distribution which means kurtusis is near to 0.
In mesokurtic distribution are moderate in breadth and curves are medium peaked height.
LEPTOKURTIC (KURTOSIS > 3)
• Leptokurtic is having very long and skinny taiłs, which means there are more chances of oulies
• Positive values of kurtosis indicate that distribution is peaked and pessesses thick. tails.
• An extreme positive kurtosis indicates a distribution where more of the numbers are located in the
tails of the distribution instead of around the mean.
PLATYKURTIC (KURTOSIS < 3)
Platykurtic having a lower tail and stretched around center tails means most of the data points are present in high proximity with mean.
• A platykurtic distribution is latter (less peaked) when compared with the normal distribution.
KURTOSIS
• Kurtosis is a statistical measure, whether the data is heavy-taíled or light-tailed
in a normal distribution.
• Kurtosis tell us about the peakedness or flaterness of the distribution Kurtosis is
basically statistical measure that helps to identify the data around.
Types of excess kurtosis
[Link] or heavy-tailed distribution (kurtosis more than normal distribution)
[Link] (kurtosis same as the normal distribution).
[Link] or short-tailed distribution (kurtosis less than normal distribution).
CALCULATE THE SKEWNESS COEFFICIENT OF THE SAMPLE
Pearson's first coefficient= Mean Mode
Pearson's second coefficient Standard Deviation
It truly scales the value down to a limited range of-1 to +1.
3(Mean Median)
Standard Deviation
Mean- Mode 3 (Mean Median)
• If the skewness is between -0.5 & 0.5. the data area nearly symmetrical
• If the skewness is between -1& -05 (negative skewed) or between 0.5 & l(positive skewed), the
data are slightly skewed.
• If the skewness is lower than-l (negative skewed} or greater than I (positive skewed), the data
are extremely skewed.
SKEWNESS
Skewness is a statistical measure used to describe the asymmetry of a probability distribution around its mean. It indicates whether the data
points in a distribution are concentrated more on one side than the other. A distribution can be positively skewed, negatively skewed, or
symmetric. Positive skewness means the tail on the right side of the distribution is longer or fatter than the left side, while negative skewness
means the opposite. If a distribution is symmetric, it has no skewness. Skewness is commonly used in finance, economics, and other fields to
understand the shape and behavior of data distributions.
TYPES OF SKEWNESS
-Positive skewed or right-skewed
- Negative skewed or left-skewed
-Symmetric Skewness
There are three main types of skewness:
1. Positive Skewness: Also known as right-skewed distribution, it occurs when the tail on the right side of the distribution is longer or fatter than
the left side. This indicates that there are more extreme values on the right side of the distribution.
2. Negative Skewness: Also known as left-skewed distribution, it occurs when the tail on the left side of the distribution is longer or fatter than
the right side. This indicates that there are more extreme values on the left side of the distribution.
3. Symmetric Skewness: This occurs when the distribution is perfectly symmetric, meaning that it has equal probabilities for values on both sides
of the mean. In this case, there is no skewness present in the distribution.
NEGATIVE SKEWED OR LEFT-SKEWED
• In which ore values are concentrated on the right side (tail)
of the distribution graph while the left tail of the distribution graph is longer.
The mean of the data is less than the median (a large number of data-pushed on the left-hand side)
frequency
• The mean, median, and mode of the distribution are negative rather than positive or zero.
(mode-median-mean)
POSITIVE SKEWED OR RIGHT-SKEWED
• In which most values are clusterd around the left tailed the distribution while the right tail of the distribution is longer.
• In positively skewed, the mean of the data is greater
than the median (a large number of data-pushed on the right-hand side),
• The mean, median, and mode of the distribution are
positive rather than negative or zero.
mode -median- mean
(positive direction)
Notes: Ds.i
i
Notes from Ds and Nev.