Quantitative Methods in IT Course
Quantitative Methods in IT Course
IT 206
Quantitative Methods
This course introduces the basic concepts of Statistics and use of quantitative
methods in research. Emphasis will be on achieving an understanding associated
mathematical tools and statistical research techniques that is relevant in solving
problems and making quality decisions across Information Technology discipline.
Students will gain experience on problem solving manually and using any statistical
tools on computers.
3 units
54 hours(lec)
IT 106
Danna Calas-Remojo
Contact Number: 09399231553
Email Address: [Link]@[Link]
Facebook/Messenger: Chikay Remojo
SLSU will:
The purpose of this course is to introduce you to the subject of statistics as a science
of data. There is data abound in this information age; how to extract useful
knowledge and gain a sound understanding of complex data sets has been more of
a challenge. In this course, we will focus on the fundamentals of statistics, which
may be broadly described as the techniques to collect, clarify, summarize, organize,
analyze, and interpret numerical information.
This course will begin with a brief overview of the discipline of statistics and will
then quickly focus on descriptive statistics, introducing graphical methods of
describing data. On the side of inference, we will focus on both estimation and
hypothesis testing issues. We will also examine statistical techniques to study the
relationship between two or more variables. By the end of this course, you should
gain a sound understanding of what statistics represent, how to use statistics to
organize and display data, and how to draw valid inferences based on data by using
appropriate statistical tools.
It will take you 54 hours(lec) to finish the four (4) modules. In each module you will
answer the Pretest to test your pre-existing knowledge of the subject, after which
you may proceed with the lessons with activities, assessments, lab exercises and
Post Test at the end of each module.
Lecture:
Performance Task (Assessment, Activities, 60%
Assignment, Post Test)
Output (Project/IDP) 40%
100%
Laboratory:
Module 2
Overview
Do the following:
In the activity provided, were you able to describe the data? What can you say
about the data presented? By using charts and tables you can easily describe and
analyze the collected sets of data. There are different tools and software for you
to choose to help you do the task and one those is descriptive statistics. Lets us
see in the following discussions what appropriate statistical tool you can use to
present and interpret your data sets.
Descriptive Statistics
- is a collection of methods for summarizing data and describe data (e.g.,
mean, median, mode, range, variance, graphs).
Frequency Distribution
- is a table or graph that organize and present the count of the number of times a
certain value or class of values occurs in a sample or population.
Frequency Tables
1. Absolute frequency (or just frequency): This tells you how many times a
particular category in your variable occurs. This is a tally, count, or
frequency of occurrence of each individual category/value in the table.
Example:
97 90 86 83 84 78 73 73 69
65 98 90 88 83 81 79 78 72
69 60 93 98 85 82 80 78 77
71 68 59 91 89 84 82 80 77
75 70 62 55 91 89 84 82 78
77 72 70 63 54 65 89 90 81
Steps:
a. Arrange data in ascending or descending order
54 55 59 60 62 63 65 65 68
69 69 70 70 71 72 72 73 75
77 77 77 77 78 78 78 78 79
80 80 81 81 82 82 82 83 83
84 84 84 85 86 88 89 89 89
90 90 90 91 91 93 97 98 98
k= 1 + 3.322 log n
where:
n= number of scores or values
k = 1 + 3.322 log 54
k = 6.76 or 7
c = R/k
where:
r= range
k= number of classes
c = R/k
c = 44/ 6.76
= 6.51 or 7
e. Fill in the lower limit- upper limit (class width) in the interval column
starting from the lowest score. There should be no overlapping of
classes and all class should have the same width. The Class Boundary is
the half points that separate the classes.
f. Tally all the scores and write the frequency of scores in each class.
g. For less than cumulative frequency (<cf) start from the lowest group of
frequency then add the frequency of each class for the succeeding
classes.
h. For greater than cumulative frequency (>cf) start from the highest group
of frequency then minus the frequency of each class for the succeeding
classes.
Frequency Distribution
2. Relative frequency (or percent): This tells you the percentage of each
category/value relative to the total number of cases.
Steps:
1. Count the total number of items in each category.
2. Divide the count (the frequency) by the total number. For example,
4/54 = .074 or 7.4%
- There are three common ways of describing the center of a set of numbers-
mean, the median and the mode
Mean
- To calculate the mean, add all the data values together and divide by the
number of values.
X =ƩX/n
Where:
X- individual score
X- mean of all scores
n- number of scores
Based on the result, it shows that the average number of nuclear power in
various country in 1989 is 38.4
Example:
Solution:
1. Find the midpoint (x) of each class by adding the upper limit and
lower limit of the class then divide it by two
X= (UL + LL)/2
Median
- Half the scores are above the median and half are below the median.
Finding median:
1. Arrange the scores into ascending order
2. If there is an odd number of numbers, the median is the middle
number
Example:
Data:
261, 160, 259, 223, 169, 127, 221, 190,224, 228, 229, 294, 204,
177, 199, 212,186, 207, 192, 241, 162, 249, 206, 210,200, 213,
185, 171, 189, 159
Sort:
127 159 160 162 169 171 177 185 186 189 190 192 199 200 204 206
207 210 212 213 221 223 224 228 229 241 249 259 261 294
Mode
- Mode is the most frequently occurring score in a distribution/data set.
Example:
15 18 19 20 20 20 21 23 23 24 24
Note:
The best one to use in a given situation depends on the type of variable
given and the interest of the researcher.
Example:
Cat 5
Dog 4
Rabbit 3
Goldfish 1
Bird 4
If the interest of the researcher is the common pets owned by the students,
then the central tendency to use is Mode.
Measure of Variability/Dispersion
- Measures of variability is defined as the dispersion (or deviation) away from
the mean for each variable.
Example:
Let’s consider the two distributions given below. They represent the
marks of a group of thirty students on two tests.
Marks on Test A
Range= 70-45 = 25
Marks on Test B
Range= 65- 45= 20
As computed, it is shown that the range in Test A is greater than the range
in Test B. Therefore, the marks on test A are more spread out than the marks
on test B
2. Standard deviation: It is the average distance of each value away from the
sample mean. The larger the standard deviation, the farther away the
values are from the mean; the smaller the standard deviation the closer,
the values are to the mean.
56 48 63 60 51 52
1. Compute Mean.
Mean = (56+48+63+60+51+52)/6= 55
2. Find how much observed values deviate from the mean by subtracting
the mean from value.
x (x-mean)
56 1
48 -7
63 8
60 5
51 -4
52 -3
3. Take the square of the deviations to get rid of the minus signs.
x (x-mean) (x-mean)2
56 1 1
48 -7 49
63 8 64
60 5 25
51 -4 16
52 -3 9
Formula:
where:
X- individual score
X- mean of sample scores
μ -mean of population scores
N- population size
n- sample size
Population Variance:
σ 2 = Ʃ ( x- μ )2 /N
= (1 + 49 + 64 + 25 + 16 + 9)/6
= 27.33.
Standard Deviation
- The standard deviation is the square root of the variance, i.e. measure of
spread in the original units. Often use Greek symbol sigma, σ for the
population standard deviation and s for the sample standard deviation.
Formula:
σ= Ʃ ( x- μ )2 /N
= (1 + 49 + 64 + 25 + 16 + 9)/(6)
= 27.33
= 5.228
Or
σ2 = 27.33
σ = 27.33
σ = 5.228
Quartile
- This measure divides the observation in four equal parts. The Q1 is the
middle point between the smallest value and the center value (median) also
called Q2. The Q2 is called the median, Q3 is the middle value between the
median and the highest value of the data set.
Example:
The median of the lower half of a set of data is the lower quartile ( LQ) or Q1 .
The median of the upper half of a set of data is the upper quartile ( UQ )
or Q3 .
If the data is even always compute the median of the values to get the
quartile values.
Example:
The upper and lower quartiles can be used to find another measure of
variation call the interquartile range .
The interquartile range or IQR is the range of the middle half of a set of
data. It is the difference between the upper quartile and the lower quartile.
In the above example, the lower quartile is 52 and the upper quartile is 58 .
Percentile
Example: If you know that your score is in the 90th percentile, that means
you scored better than 90% of people who took the test.
Example:
Consider the sugar levels of 20 students and look for the 40th percentile and
28th percentile.
80 85 88 90 90 91 91 98 100 102
109 111 120 122 123 124 125 130 140 201
L=(K/100) * N
Note: If L is whole number, kth is midway between L and the next value. If L
is not a whole number, round it up to the next integer and that value is the
kth percentile.
40th percentile:
L=(40/100)*20=8
Thus, the 40th percentile is the value between 8th and 9th observation and it
is 99
28th percentile:
L=(28/100)*20=5.6
Thus, the 28th percentile is the 6th item in the observation which is 91
Z-score
Formula:
z = (X - μ) / σ
where:
z - z-score
X- value of the element
μ - mean of the population
σ - standard deviation
Example:
Solution:
z = (X - μ) / σ
z= (65-60)/5 = 1
The result shows that 65 is one standard deviation higher than the mean
Measure of Shape
- is a measure of the asymmetry of the probability distribution of a random
variable about its mean
Symmetrical Distribution
Asymmetric Distribution
- In Asymmetric Distribution, the two sides do not mirror each other because
the collected sample data is systematically biased.
- Plotting such distribution on a histogram, the tail on the right side of the
histogram is longer than the left side with the mean being on the right and
median on the left and is the reason that Positively Skewed Distribution is
also known as Right Skewed Distribution.
Mode<median<mean
- Plotting such a distribution on a histogram, the tail on the left side of the
histogram is longer than the right side with the mean being on the left and
median on the right and is the reason that Negatively Skewed Distribution
is also known as Left Skewed Distribution.
mean<median<mode
Steps:
1. Find the median of the class as x
2. Find the mean for grouped data using the formula x = Ʃfx/N
s= Ʃ ( x- x )2 /N-1
= 35.6/9
=1.9889
5. Calculate Skewness
Skewness= Ʃ ( x – x )3
(n-1) s3
= 34.56
9 (1.9889)3
= 0.4881
Rule:
Since the skewness value in the example is 0.4881 and it is less than 1,
therefore the distribution of data is negatively skewed (left)
Kurtosis
Leptokurtic
Platykurtic
- The data points are highly dispersed along the X-axis that results in thinner
tails when compared to a normal distribution and has very few values
clustered around the center (mean).
Mesokurtic
- Here the tails of the distribution are neither too thin nor they are too thick
and also the scores are equally divided with scores neither being clustered
around the center nor being too scattered.
Kurtosis= Ʃ ( x – x )4
(n-1) s4
= 299.79
9 (1.9889)4
= 299.79
9 (15.6464)
= 2.1289
Rule:
Since the kurtosis value in the example is 2.1289 which is greater than 1,
therefore the shape of distribution is Leptokurtic
1. BAR GRAPH- is a graphical display of data using bars and can be plotted
vertically or horizontally.
Source: [Link]
Source: [Link]
3. LINE GRAPH- is a type of chart that uses lines to connect individual data
points displaying quantitative values over a specified time interval.
Source: [Link]
Source: [Link]
Source: [Link]
IT 206 QUANTITATIVE METHODS MODULE DANNA CALAS-REMOJO
29
Steps:
1. Type your data into a worksheet. Make sure you put your summarized data
into columns (category, frequency and relative frequency). Use column
headers and note that column headers will become the labels on the graph.
2. For ungroup data, simply highlight the category and frequency column and go
to insert menu then select the type of chart you want to use.
3. For grouped data, organized your data in the worksheet and to create the
chart, follow instructions provided from
[Link]
Group activity: