0% found this document useful (0 votes)
9 views34 pages

Data Management in Modern Mathematics

This instructional module from Nueva Vizcaya State University covers data management, emphasizing the importance of accurate data collection, organization, and interpretation. It outlines desired learning outcomes, course content, and various statistical methods, including measures of central tendency and dispersion, as well as data presentation techniques. The module aims to equip students with the skills to effectively manage and analyze data in their respective fields.

Uploaded by

renzemaria936
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views34 pages

Data Management in Modern Mathematics

This instructional module from Nueva Vizcaya State University covers data management, emphasizing the importance of accurate data collection, organization, and interpretation. It outlines desired learning outcomes, course content, and various statistical methods, including measures of central tendency and dispersion, as well as data presentation techniques. The module aims to equip students with the skills to effectively manage and analyze data in their respective fields.

Uploaded by

renzemaria936
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Republic of the Philippines

NUEVA VIZCAYA STATE UNIVERSITY


Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

College : Engineering
Campus: Bambang

DEGREE PROGRAM BSCpE COURSE NO. GE Math


SPECIALIZATION Computer COURSE TITLE Mathematics in the Modern World
YEAR LEVEL 1st Year TIME FRAME Hrs WK NO. 10-13 IM NO. 4

I. UNIT TITLE/CHAPTER TITLE


Data Management

II. LESSON OVERVIEW


Data management is the process of ingesting, storing, organizing and maintaining
the data created and collected by an organization. The data management process includes a
combination of different functions that collectively aim to make sure that the data in corporate systems is
accurate, available and accessible. Data management further explains what it is and provides insight on
the individual disciplines it includes, best practices for managing data, challenges that organizations face
and the business benefits of a successful data management strategy. You'll also find an overview of data
management tools and techniques.
The use of statistical methods in manufacturing, development of food products, computer
software, energy sources, pharmaceuticals, and many other areas involves the gathering of information
or scientific data. Of course, the gathering of data is nothing new. It has been done for well over a
thousand years. Data have been collected, summarized, reported, and stored for perusal. However, there
is a profound distinction between collection of scientific information and inferential statistics. It is the latter
that has received rightful attention in recent decades.

III. DESIRED LEARNING OUTCOMES

1. Use a variety of statistical tools to process and manage numerical data;


2. Use the methods of linear regression and correlations to predict the value of a variable given certain
conditions; and
3. Advocate the use of statistical data in making important decisions.
4. Enumerate and explain classifications of Random Variables
5. Explain properties of probability function.
6. Explain cumulative distribution function.
7. Explain counting rules useful in probability

IV. COURSE CONTENT

A. Gathering, Organizing, Representing and Interpreting Data


1. Data Gathering
2. Data Organization and Presentation
3. Data Analysis and Presentation
B. Measures of Central Tendency
1. Mean
2. Median
3. Mode
C. Measures of Dispersion
1. Range
2. Mean Absolute Deviation or Variance
3. Standard Deviation

NVSU-FR-ICD-05-00 (081220) Page 1 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

4. Coefficient of Variation
5. Skewness
6. Coefficient of Kurtosis
D. Measures of Relative Position
1. Percentiles
2. Deciles
3. Quartiles
E. Probabilities and Normal Distribution

V. LESSON CONTENT

Definition of Terms`

1. Statistics is a science that deals with the methods of collecting, organizing, summarizing and
interpreting data in order to draw valid conclusions from them.
2. Descriptive Statistics: concerned with the collection and presentation of data and the description of
some of their features to yield meaningful information without attempting to draw any inferences from
them.
3. Inductive or inferential statistics: concerned with the development and used of mathematical tools
to go beyond data presentation and make forecasts and inferences.
4. Population: consist of all the individuals or objects in a group under study.
5. Sample: a sub collection of items drawn from a population under study.
6. Variable: characteristics that is being studied. It may be qualitative or quantitative.
7. Data: facts and statistical collected together for reference or analysis.
8. Raw data: collected data that have been organized numerically. An arrangement of raw data in
ascending or descending order or magnitude is an array.
9. Frequency: number of times a value appears in the listings.
10. Cumulative frequency: the total frequency of all values either “ less than” or more than” any class
boundary.

Statistics deals also with the development of techniques for collection data. Data should be properly
collected so that an investigator may be able to answer the question under consideration with a
reasonable degree of confidence.
The simplest method for ensuring a representative selection of samples is to take a simple random
random sample.
In stratified sample method this involves taking sample fom each population unit in non-
overlapping groups.

Different types of Data:

Primary Data: Data that has been collected from the first hand experiences is known as primary data. It
has more reliable, authentic and not been published anywhere.
Secondary Data: Data that those have already collected by others.
These are usually maybe available in the published or unpublished form. When it is not
possible to collect the data by primary method, the investigator go for secondary method.
This data collected for some purpose other than the problem at hand.

NVSU-FR-ICD-05-00 (081220) Page 2 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

2. Factors to be considered in collection of Data.

1. Objects and scope of the inquiry


2. Sources of Information
3. Quantitative expression
4. Techniques of data collection
5. Unit of collection.
6. Sources of Data

Gathering, Organizing, Representing and Interpreting Data


An investigation should always be based on accurtae data which requires good management.
Correct methods of collecting data, right way of organizing them and goes data presentation will result to
a precise analysis and interpretation.
All the various steps of data processing should be planned when the study is designed before any
data are collected.

NVSU-FR-ICD-05-00 (081220) Page 3 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

1. Data Gathering

There are different methods used in gathering or collecting data. These are:

1. Direct or interview method is a person to person encounter between the source of information,
the interviewee, and the one who gathers information, the interviewer. Interview can be done
personal, through phone or interview access.
2. Indirect or questionaire method is the technique in which questionaire is used to elicit the
information or data needed.
3. Registration method obtains data from the records of government agency authorized by law to
keep such data or information and made these available to researchers.
4. Observation method is a technique in which data particularly thoses pertaining tot he behaviors
of individuals during the given situation are best obtained through observations.
5. Experimental method is a system used to gather data from the results of performed series of
experiments on some controlled and experimental variables. This is commonly used in scientific
inquiries.

2. Data Organizations and Presentation

Data collected or obtained from whatever manner are called raw data. Data collected can be classified
according to the scale of measurement used. There are four levels of measurements from lowest to highest
scale: the nominal, ordinal, interval and ratio scales.

a) Nominal scale assigns names or labels to observation in purely arbitrary sequence. The labels are used
to classify the respondents or objects without ordering. For instance, if we need to classify the
respondents’ preference on cellphone brands, such as apple, Samsung, Lg, one plus, oppo etc. we
measure on the nominal scale and data gathered is a nominal data.
b) Ordinal scale assigns numbers or labels to observations with implied ordering. Ranking the respondents
preferences means measuring responses in the ordinal scale and the data obtained is called ordinal data.
c) Interval scale assigns real numbers to observation to reflect distance between rank positions of the
respondents or objects in equal units. This scale gives the distance between any two numbers of known
sizes, has zero point and has a unit measurement. The data collected can be manipulated algebraically
by addition or subtraction but not division or multiplication. Examples are expenses, distance, weight etc.
d) Ratio scale assigns number to observations to reflect the existence of true absolute zero point origin as
it’s origin. The ratio of two scale point is independent of the unit of measurement. The data collected has
all properties of an interval data and manipulated algebraically by multiplication and division. Examples
are on birth rate, unemployment rate etc.
Take note that the data set obtained using interval and ratio scales is called measured data.
An effective presentation of data is necessary in any investigation. Data, which have been
collected and organized well, but not presented properly and clearly would become of little use. Data
presentation should give a clear picture of various relationship among the data presented. Proper
presentation of gathered data is important in order to decide what tool or tools of analysis should be
employed to come up with an intelligent judgement.

There are different ways of forms to present data. These are: textual, tabular and graphical.
1. Textual form make use of words, sentences and paragraph in presentation. It is commonly used
when there are only few numerical data to be enumerated or to be compared with other data.
NVSU-FR-ICD-05-00 (081220) Page 4 of 34
Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

2. Tabular form is a systematic presentation of data rows and columns. It is used when related
numerical facts need to be classified in arrays.
3. Graphical form shows numerical values or relationships in pictorial form. It makes used of graphs,
symbols or visual aids.

Tabular Presentations must be simple, focus the reader’s attention on the data rather than on the
form and it should make the meanings and significance of information being presented clear.
a. It should be simple
b. It should focus the reader’s attention on the data rather than on the form.
c. It should make the meanings and significance of information being presented clearly.

Statistical Tables should have the following parts:


1. Heading which shows the table number, title and head note.
The title is brief statement of nature, classification and time reference of the information
presented and the area to which the statistics refer.
The head note is a statement enclosed in brackets between the title and top rule of the
table that provides additional information.
2. Box head is the portion that contains the column heads, which describe the data in each
column.
3. Stub is the first column on the left of the table, which describes the data on the given row.
4. Footnote is a statement inserted at the bottom of the table.
5. Source note is exact citation of the source of data which is usually include to acknowledge
the origin of the data.

NVSU-FR-ICD-05-00 (081220) Page 5 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Graphical Presentation

Good chart should possess the following properties:


1. Accurate the dimensional aspects should reflect the highest degree accuracy possible within the
practical limits imposed by expert draftsman or the electronic computer being used. It should not
be deceptive, distorted, or misleading or in any way susceptible to wrong interpretation as a result
of inaccurate or careless construction.
2. Simple the basic design should be simple and straight forward and not loaded with irrelevant or
trivial symbols and ornamentation.
3. Clear it should be easily read understood. There should be forceful and unmistakable focus of
the message that the graph is trying to communicate and there should be a truthful and
unambiguous representation of the facts and that the message it conveys is meaningful.
4. Attractive it is designed and constructed to attract and hold the attention by holding a neat,
dignified and professional appearance. It should be stylish.

Different types of graphs can be used in data presentation based on purpose. These are:

Line Graph is used when ( 1 ) data cover a long period of time ( 2 ) several series are compared (3)
movements are to be emphasized ( 4 ) trends are to be established and ( 5 ) estimates are to be
forecasted.

Bar Graph is used when numerical values of an item over a period of time are compared. It consists
of regular bars where the height of bars represents quantity or frequency for each category.

NVSU-FR-ICD-05-00 (081220) Page 6 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Pie graph is used to show percentage or the composition by parts of a whole.

Pictograph or pictogram is used to immediately suggest the nature of data.

Organizing collected numerical data can be done in two ways

1. Array is an arrangement of the numerical data/values according to order of magnitude either


ascending or descending.
2. Frequency distribution table is condensed version of an array. It categorizes the numerical
data into intervals or classes. It has the following parts:
Classes are mutually exclusive categories defining the lower limit and the upper limit
with equal intervals.
Class frequency is the number of observations in each class.
Class mark or class midpoint is used in computing the mean and some measures of
variability.
Cumulative Frequency tells the sum of frequencies in a particular class of interest.
tells the percentage of observations in a particular class of interest.

Steps in Constructing a Frequency Distribution with Equal Class Size

1. Determine the range R of the numerical data.


R= / Highest Value- Lowest Value
2. Determine the number of classes K to which the data are to be grouped using the Sturges’
Approximation:
K= 1+ 3.322 Log N
where N= total number of values to be grouped
3. Determine the class size C
C= R/K
4. Determine the lower limit of the first class.
Note: There is no fixed rule in determining the lower limit of the first class.
For the purpose of uniform result, the lowest value in data set should be the lower l9mit of
the first class.
5. Construct the class intervals and determine the class frequencies.
NVSU-FR-ICD-05-00 (081220) Page 7 of 34
Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Remarks:
Sturges’ Approximation
1. Sturges’ Approximation is just a guide and flexible rule.
2. The number of classes should be large enough to demonstrate the major characteristics
of the data yet not so large as to result in losing the advantage of summarizing raw data.
For instance, where the highest observed value fails to be included in the last class
constructed, the number of classes should not be increased just to accommodate the
highest value increase in class size.
3. The number of classes is usually taken between 5 to 20 depending nature of the data
without using the Sturges’ Approximation.
4. Class intervals are chosen so that the class marks coincide with actually observed data.

Graphical Presentation of Frequency Distribution with Equal Class Size

Steps in Constructing Frequency Charts


1. Label either class limits or class marks along the horizontal axis.
2. Plot the frequency of each class along the vertical axis above the class mark of the corresponding
class.
3. The vertical scale must always include zero.
4. The horizontal scale must include only the range of the observed data and one extra interval at
each end.
5. The vertical axis height should be approximately ¾ the length of the horizontal axis.

Frequency Histogram a set of vertical bars whose areas are proportional to the frequencies
presented.

Note: The length of the bars is equal to the class size and the height is numerically equal to the class
frequency.

Frequency Polygon is a line chart plotted along the same scale as the histogram. The class frequency
is plotted against the class mark.

NVSU-FR-ICD-05-00 (081220) Page 8 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

3. Data Analysis and Interpretation

Data analysis and interpretation is the process of making sense of numerical data that has been
collected, analyzed and presented. A common method of describing the characteristics of individual
objects or group of individuals under study is known as descriptive statistics, while the analyzing and
interpreting data is known as inferential statistics.
Descriptive statistics give a single value while represents the set values. There are three
methods of describing a set values: the measures of central tendency, measures of dispersion and
measure of skewness and kurtosis. Measures of central tendency refer to a value where the set of
values differ from each other: while skewness and kurtosis measures the symmetry and flatness/
peakedness of the distribution.
Inferential statistics are techniques wherein samples can be used to make generalizations about
the population from which the samples were drawn. It is important that the sample accurately
represents the population. The process of achieving this is called sampling. There are two methods of
inferential statistics: the estimation of parameters and hypothesis testing.

A. Measures of Central Tendency

Methods of central tendency are measures indicating the center of a set data which arranged in
order of magnitude. It is described as the point about which the scores tend to cluster, hence, regarded
as a sort of average in the series. It is the center of the concentration of the scores. It is a single
number which describes the totality of the set of data collected. It refers to the parameters of the
sample.

There are three measures of central tendency commonly used. These are mean, median and
mode.

Mean or arithmetic mean ( average ) is the most popular and well known measure of central
tendency. It can be used with both discrete and continuous data ( although it is used most often with
continuous data.

Median is defined as the middle value when a set of observed values have been arranged in either
ascending ( from lowest to highest ) or descending ( from highest to lowest ) order of magnitude. The
median is the centermost array into two equal parts, that is 50% of the total number of observation is
less than the median value while the other is 50% is greater than the median value.

Mode is the most frequent score in the data set. It is sometimes considered as the most popular option.

B. Measures of Dispersion
Range – It is the simplest measures of dispersion. It is the difference between the highest and the
lowest point.
R = HV – LV
Mean absolute deviation – also known as variance. It is the simplest method of taking into account
the variations or the spread ability of all items into a series from the point of central tendency.
NVSU-FR-ICD-05-00 (081220) Page 9 of 34
Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

ð2 =Ʃ( X1 - µ)2/N
Standard deviation – is based on the deviations of all the scores in a [Link] is always computed
from the mean. The SD is defined as the positive square root of the variance.
S = √𝑺2

The absolute and mean absolute deviation show the amount of deviation (variation) that occurs around
the mean score. To find the total variability in our group of data, we simply add up the deviation of each
score from the mean. The average deviation of a score can then be calculated by dividing this total by the
number of scores. How we calculate the deviation of a score from the mean depends on our choice of
statistic, whether we use absolute deviation, variance or standard deviation.
Perhaps the simplest way of calculating the deviation of a score from the mean is to take each score
and minus the mean score. For example, the mean score for the group of 100 students we used earlier
was 58.75 out of 100. Therefore, if we took a student that scored 60 out of 100, the deviation of a score
from the mean is 60 - 58.75 = 1.25. It is important to note that scores above the mean have positive
deviations (as demonstrated above), whilst scores below the mean will have negative deviations.
To find out the total variability in our data set, we would perform this calculation for all of the 100
students' scores. However, the problem is that because we have both positive and minus signs, when we
add up all of these deviations, they cancel each other out, giving us a total deviation of zero. Since we are
only interested in the deviations of the scores and not whether they are above or below the mean score, we
can ignore the minus sign and take only the absolute value, giving us the absolute deviation. Adding up
all of these absolute deviations and dividing them by the total number of scores then gives us the mean
absolute deviation. Therefore, for our 100 students the mean absolute deviation is 12.81.

COEFFICIENT OF VARIATION

Coefficient of Variation – also known as relative dispersion, hence it can be used to compare
variability of two or more groups of data measured in the same or different units.

SKEWNESS AND KURTOSIS

Skewness – is a measure or a criterion on how symmetric the distribution of data is from the mean.
Positive skewness – indicates a distribution with an asymmetric tail extending toward the right side of the
distribution.
Negative skewness – indicates a distribution with an asymmetric tail extending toward the left.

Using measures of central tendency


1. If mean = median = mode, the skewness is zero. ( symmetrical )
2. If mean > median > mode, the skewness is positive.
3. If mean < median < mode, the skewness is negative.

Kurtosis – measures the flatness and peakedness of the distribution of a given data set. It also measures
the degree of departure from the normal distribution.

1. Leptokurtik – a distribution which is more peaked than the normal distribution.


2. Platykurtic - a distribution which is flatter than the normal distribution.
3. Mesokurtic – it is normal in shaped.

NVSU-FR-ICD-05-00 (081220) Page 10 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

N = 10

NVSU-FR-ICD-05-00 (081220) Page 11 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

NVSU-FR-ICD-05-00 (081220) Page 12 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

C. MEASURES OF RELATIVE POSITION

NVSU-FR-ICD-05-00 (081220) Page 13 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

1. RANGE R = HV – LV
2. PERCENTILE DEVIATION PD = P90 – P10
3. DECILE DEVIATION DD = D90 – D10
4. INTERQUARTILE RANGE IR = Q9 – Q1
5. SEMI-INTERQUARTILE RANGE QD = Q9 – Q1 / 2 OR IR / 2
6. AVERAGE DEVIATION AD = Ʃf IxI / N
7. STANDARD DEVIATION SD

NVSU-FR-ICD-05-00 (081220) Page 14 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

NVSU-FR-ICD-05-00 (081220) Page 15 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

NVSU-FR-ICD-05-00 (081220) Page 16 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

EXAMPLE
48 85 91 54 62 72 68 70 94 98 62 76 99 64
71 49 89 68 98 66 96 55 77 57 91 79 53 62
59 82 93 60 79 90 59 43 52 61 88 73 51 69
100 93 61 70 92 46 73 83

MASTER SHEET
Is a tool or a device used in arranging test score in statistical data.

NVSU-FR-ICD-05-00 (081220) Page 17 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

EXAMPLE

HV = 100 LV = 43

HV – LV 100 4 IS CONSTANT horizontal lines


_43 13 IS CONSTANT vertical lines
6 + 4 = 10

0 1 2 3 4 5 6 7 8 9 TOTAL
10 / 1
9 / // / // / / // / 11
8 / / / / / 5
7 // / / // / / // 10
6 / // /// / / // / 11
5 / / / / / / // 8
4 / / / / 4
TOTAL 5 6 7 7 3 2 4 2 6 8 50

a. use the master sheet to tally scores and rank


b. solve for the mean, median, mode
-ungrouped
-grouped

RANKING

SERIAL DATA RANK SERIAL DATA RANK


NUMBER NUMBER
1 100 1 26 70 26.5
2 99 2 27 70 26.5
3 98 3.5 28 69 28
4 98 3.5 29 68 29.5
5 96 5 30 68 29.5
6 94 6 31 66 31
7 93 7.5 32 64 32
8 93 7.5 33 62 34
9 92 9 34 62 34
10 91 10.5 35 62 34
11 91 10.5 36 61 36.5
12 90 12 37 61 36.5
13 89 13 38 60 38
14 88 14 39 59 39.5
15 85 15 40 59 39.5
16 83 16 41 57 41
17 82 17 42 55 42
18 79 18.5 43 54 43
19 79 18.5 44 53 44
20 77 20 45 52 45
21 76 21 46 51 46
22 73 22.5 47 49 47
23 73 22.5 48 48 48
24 72 24 49 46 49
25 71 25 50 43 50

NVSU-FR-ICD-05-00 (081220) Page 18 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Solving for the mean, median and mode

GROUPED DATA

Data Tally f Exact limits Mid <cf >cf <RCF >RCF RF %RF
UL LL point
100-104 I 1 99.5-104.5 102 50 1 1.000 0.0200 0.0200 2.0
95-99 IIII 4 94.5-99.5 97 49 5 0.980 0.1000 0.0800 8.0
90-94 IIIII-II 7 89.5-94.5 92 45 12 0.900 0.2400 0.1400 14.0
85-89 III 3 84.5-89.5 87 38 15 0.760 0.3000 0.060 6.0
80-84 II 2
75-79 IIII 4
70-74 IIIII-I 6
65-69 IIII 4
60-64 IIIII-II 7
55-59 IIII 4
50-54 IIII 4
45-49 III 3
40-44 I 1
N=50 ƩRF=1.00
Ʃ%RF=100%

<RCF = <cf / N RF = f / n

>RCF = >cf / N

PROBABILITY DISTRIBUTION

A probability distribution is a function that describes the likelihood of obtaining the possible
values that a random variable can assume. In other words, the values of the variable vary based on
the underlying probability distribution.

Basics of Probability Distributions

As a reminder, a variable or what will be called the random variable from now on, is represented
by the letter x and it represents a quantitative (numerical) variable that is measured or observed in an
experiment.

Also remember there are different types of quantitative variables, called discrete or continuous.
What is the difference between discrete and continuous data? Discrete data can only take on particular
values in a range. Continuous data can take on any value in a range. Discrete data usually arises from
counting while continuous data usually arises from measuring.

NVSU-FR-ICD-05-00 (081220) Page 19 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

1.a.

Probability
1. In the study of statistics, we are concerned basically with the presentation and interpretation
of chance outcomes that occur in a planned study or scientific investigation. For example, we
may record the number of accidents that occur monthly at the intersection of Driftwood Lane
and Royal Oak Drive, hoping to justify the installation of a traffic light; we might classify items
coming off an assembly line as “defective” or “non defective”; or we may be interested in the
volume of gas released in a chemical reaction when the concentration of an acid is varied.
Hence, the statistician is often dealing with either numerical data, representing counts or
measurements, or categorical data, which can be classified according to some criterion.
2. We shall refer to any recording of information, whether it be numerical or categorical, as an
observation. Thus, the numbers 2, 0, 1, and 2, representing the number of accidents that
occurred for each month from January through April during the past year at the intersection of
Driftwood Lane and Royal Oak Drive, constitute a set of observations. Similarly, the
categorical data N, D, N, N, and D, representing the items found to be defective or non
defective when five items are inspected, are recorded as observations.
3. Statisticians use the word experiment to describe any process that generates a set of data. A
simple example of a statistical experiment is the tossing of a coin. In this experiment, there
are only two possible outcomes, heads or tails. Another experiment might be the launching of
a missile and observing of its velocity at specified times. The opinions of voters concerning a
new sales tax can also be considered as observations of an experiment. We are particularly
interested in the observations obtained by repeating the experiment several times. In most
cases, the outcomes will depend on chance and, therefore, cannot be predicted with certainty.
If a chemist runs an analysis several times under the same conditions, he or she will obtain
different measurements, indicating an element of chance in the experimental procedure. Even
when a coin is tossed repeatedly, we cannot be certain that a given toss will result in a head.
However, we know the entire set of possibilities for each toss.

1.b. Sample Space of An Experiment

NVSU-FR-ICD-05-00 (081220) Page 20 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

An experiment in any action or process that generates data. The set of all possible outcomes of
an experiment is the sample space. Each outcome in a sample space is called an element or sample
point.

Methods of Describing Data

1. If the sample space has finite number of sample points, we may describe the set listing the
elements separated by commas and enclosed brackets.
2. If the sample space has large or infinite number of sample points, describe the set by
statement or rule.

Tree Diagram

In some experiments, it will be helpful to list the elements of s systematically by means of a tree
diagram.

Consider the experiment of examining three bulbs. We note the results of the examination by
constructing tree diagram as shown in the figure.

The branches of the tree give the distinct sample points. Starting at the top and by following the
paths of the branches, the sample space S is

S= {DDD, DDN, DND, DNN, NDD, NDN, NND, NNN}

Events

In the study of probability, we will be interested in any collection of outcomes in S rather than
individual outcomes of S. Any sub collection of outcomes in the sample space S is called event.

For instance, in examining three bulbs, A={ DDN,DND,NDD} is the event that exactly two bulbs
are defective. B={ DDD, DDN,DND,NDD} is the event that at least two bulbs are defective.

NVSU-FR-ICD-05-00 (081220) Page 21 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Some Relations from Set Theory

1. A ∪ B is the event “ either A or B or both will happen. This connotes the occurrences of event
A or B or the simultaneous occurrence of A and B
2. A ∩ B is the event both A and B will happen. This connotes the simultaneous occurrence of
A and B.
If A ∩ B= Ø, A and B are mutually exclusive events. Mutually exclusive events are events that
cannot occur simultaneously.
3. A’, the complement of a is the event “ not A” will happen. A’ is the set of all outcomes in s
that are not in A.

Example: Two bulbs are placed in two sockets inside a dark room. Use G for the bulb that will
light up, the bulb in good condition and B for the bulb which will not light up. Make a sample space
for the two light bulbs.

S={ GG, GB, BG }

A. At least one of the bulbs will light up


A= { GG, GB, BG}
B. Both are the same condition
B= { GG, BB }
C. Neither bulb will light up
C= { BB }
D. At least one of the bulbs will light up or both are the same condition
D= { GG, GB, BG, BB }
E. At least one of the bulbs will light up and both are defective
E= { }=Ø
Event D= event “A or B” = A∪B.
Event E= event “A and C”=Ø since events A and C are mutually exclusives events.

1.c. Counting Techniques

If the number of possible outcomes in an experiment is quite large, the effort of constructing the
list of outcomes becomes prohibited. By using some counting rules, it is possible to determine the number
of outcomes without listing.

1.d. Fundamental Principle ( Multiplication Rule )

The fundamental principle in counting rules which also referred to as the multiplication rule can
be stated as follows:

If an operation can be performed in n ways, and if for each of these a second operation can be
performed in n2 ways, and for each of the first two a third operation can be performed in n3 ways and so
forth, then the sequence of k operations can be performed in n=n1*n2*n3……nk ways.

NVSU-FR-ICD-05-00 (081220) Page 22 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Illustrative Problems:

2. A man wishes to travel town A to town D. There are two roads he may choose from in traveling from
town A to B, three roads that connect towns B and C and four roads that he may opt to travel from
town C to town D. In how many ways can the man travel from town A to D?

Solution:

n= n1* n2* n3 = 2(3)(4)= 24 ways

3. From the digits 2,4,6,8 and 9


a. How many distinct three digit numbers can be formed?
b. How many of these are even?

Solution:

a. 𝑛1 = number of choices for the ones place value= 5 digits as choices


n2 = number of choices for the tens place value = 4 digits
n3 = number of choices for the hundred place value = 3 digits
n= 5 ( 4 ) ( 3 )= 60 distinct three- digit numbers
b. 𝑛1 = number of choices for the ones place values = 2 digits
n2 = number of choices for the tens place value = 4 digits
n3 = number of choices for the hundreds place value = 3 digits
n = 2 ( 4 ) ( 3 ) = 24 three digit even numbers

Alternative Solutions for question a

Since the digits should be distinct, there, there are 5 digits to be arranged by 3s, such number of
arrangement is

NVSU-FR-ICD-05-00 (081220) Page 23 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

4. How many distinct permutations are there in the word MILLENNIUM?


Solutions:
There are 2M’s, 2L’s, 2 I’s and 2 N’s
𝟏𝟎!
P= = 226,800
𝟐!𝟐!𝟐!𝟐!
5. a. In how many ways can 4 letters a, b, c, d, be arranged in a circle? b. How many arrangements
are there if a and b must always be together?

Solutions:

a. P = (4-1 ) ! = 3!= 6 ways


b. So that a and b always together, arrange only 3 positions in a circle, thus
n = ( 3-1 ) ! = 2! = 2
n2 = 2!= number of ways letters a and b arranged
n = 2 ( 2 )= 4 arrangements
6. From a box containing 4 defective and 5 non defective items, how many samples of size 3 are
possible,
i. With restrictions

Solution:

THE NORMAL DISTRIBUTION

A normal distribution, sometimes called the bell curve, is a distribution that occurs naturally in many situations.

The empirical rule tells you what percentage of your data falls within a certain number of standard deviations from the mean:
• 68% of the data falls within one standard deviation of the mean.

• 95% of the data falls within two standard deviations of the mean.

• 99.7% of the data falls within three standard deviations of the mean.

NVSU-FR-ICD-05-00 (081220) Page 24 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

The standard deviation controls the spread of the distribution. A smaller standard deviation indicates that the data
is tightly clustered around the mean; the normal distribution will be taller. A larger standard deviation indicates that
the data is spread out around the mean; the normal distribution will be flatter and wider.

Properties of a normal distribution


1. The mean, mode and median are all equal.
2. The curve is symmetric at the center (i.e. around the mean, μ).
3. Exactly half of the values are to the left of center and exactly half the values are to the right.
4. The total area under the curve is 1.

Common Properties for All Forms of the Normal Distribution

Despite the different shapes, all forms of the normal distribution have the following characteristic properties.

1. They’re all symmetric. The normal distribution cannot model skewed distributions.
2. The mean, median, and mode are all equal.
3. Half of the population is less than the mean and half is greater than the mean.
4. The Empirical Rule allows you to determine the proportion of values that fall within certain distances from the
mean. More on this below!

While the normal distribution is essential in statistics, it is just one of many probability distributions, and it does not
fit all populations. To learn how to determine whether the normal distribution provides the best fit to your sample
data, read my posts about How to Identify the Distribution of Your Data and Assessing Normality: Histograms vs.
Normal Probability Plots.

Characteristics of the Normal Distribution

a. The normal distribution is a continuous distribution in which random variable X can assume value
between -∞≤≤∞.
b. The two parameters that can describe the normal distribution are mean µ and the variance δ2.
c. The normal distribution is a symmetric, bell-shaped probability distribution.
NVSU-FR-ICD-05-00 (081220) Page 25 of 34
Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

d. The total area under the normal curve and above the x-axis is one.
e. The normal curve approaches the horizontal axis asymptotically as the normal curve extends in either
direction from the mean.

LINEAR REGRESSION AND CORRELATION

A scatterplot can identify several different types of relationships between two variables.

1. A relationship has no correlation when the points on a scatterplot do not show any pattern.
2. A relationship is non-linear when the points on a scatterplot follow a pattern but not a straight line.
3. A relationship is linear when the points on a scatterplot follow a somewhat straight line pattern. This is the relationship
that we will examine.

Linear relationships can be either positive or negative. Positive relationships have points that incline upwards to
the right. As x values increase, y values increase. As x values decrease, y values decrease. For example, when
studying plants, height typically increases as diameter increases.

Figure 2. Scatterplot of height versus diameter.

Negative relationships have points that decline downward to the right. As x values increase, y values decrease.
As x values decrease, y values increase. For example, as wind speed increases, wind chill temperature decreases.

Figure 3. Scatterplot of temperature versus wind speed.

Non-linear relationships have an apparent pattern, just not linear. For example, as age increases height increases
up to a point then levels off after reaching a maximum height.

NVSU-FR-ICD-05-00 (081220) Page 26 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Figure 4. Scatterplot of height versus age.

When two variables have no relationship, there is no straight-line relationship or non-linear relationship. When one
variable changes, it does not influence the other variable.

F. Linear Correlation Coefficient

Because visual examinations are largely subjective, we need a more precise and objective measure to define the
correlation between the two variables. To quantify the strength and direction of the relationship between two
variables, we use the linear correlation coefficient:

where x̄ and sx are the sample mean and sample standard deviation of the x’s, and ȳ and sy are the mean and
standard deviation of the y’s. The sample size is n.

An alternate computation of the correlation coefficient is:

Where:

NVSU-FR-ICD-05-00 (081220) Page 27 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

The linear correlation coefficient is also referred to as Pearson’s product moment correlation coefficient in honor of
Karl Pearson, who originally developed it. This statistic numerically describes how strong the straight-line or linear
relationship is between the two variables and the direction, positive or negative.

The properties of “r”:

1. It is always between -1 and +1.


2. It is a unitless measure so “r” would be the same value whether you measured the two variables in pounds and inches or
in grams and centimeters.
3. Positive values of “r” are associated with positive relationships.
4. Negative values of “r” are associated with negative relationships.

Examples of Positive Correlation

Correlation
Is a relationship or association between variables.
Linear correlation coefficient
Denoted by ρ (rho), is a measure of the strength of the linear relationship existing between two variables,
X & Y, which is independent of their respective scale of measurement.

NVSU-FR-ICD-05-00 (081220) Page 28 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Correlation Coefficient Interpretation


-1.00 Perfect negative Correlation
-0.76 to – 0.99 Very High Negative Correlation
-0.51 to -0.75 High Negative Correlation
-0.26 to -0.50 Moderately Small Negative Correlation
-0.01 to -0.25 Very Small Negative Correlation
0.00 No Correlation
0.01 to 0.25 Very Small Positive Correlation
0.26 to 0.50 Moderately Small Positive Correlation
0.51 to 0.75 High Positive Correlation
0.76 to 0.99 Very High Positive Correlation
1.00 Perfect Positive Correlation

Analysis Regression and Correlation Simple Linear and Correlation.

This function provides simple linear regression and Pearson's correlation.

Regression parameters for a straight line model (Y = a + bx) are calculated by the least squares method
(minimization of the sum of squares of deviations from a straight line). This differentiates to the following formulae

for the slope (b) and the Y intercept (a) of the line:
Regression assumptions:
1. Y is linearly related to x or a transformation of x
2. deviations from the regression line (residuals) follow a normal distribution
3. deviations from the regression line (residuals) have uniform variance

A residual for a Y point is the difference between the observed and fitted value for that point, i.e. it is the distance
of the point from the fitted regression line. If the pattern of residuals changes along the regression line then consider
using rank methods or linear regression after an appropriate transformation of your data.
Pearson's product moment correlation coefficient (r) is given as a measure of linear association between the two

variables:
r² is the proportion of the total variance (s²) of Y that can be explained by the linear regression of Y on x. 1-r² is the
proportion that is not explained by the regression. Thus 1-r² = s²xY / s²Y.

Confidence limits are constructed for r using Fisher's z transformation. The null hypothesis that r = 0 (i.e. no
association) is evaluated using a modified t test (Armitage and Berry, 1994; Altman, 1991).

Pearson's correlation assumption:


· at least one variable must follow a normal distribution

The estimated regression line may be plotted and belts representing the standard error and confidence interval for
the population value of the slope can be displayed. These belts represent the reliability of the regression estimate,
the tighter the belt the more reliable the estimate (Gardner and Altman, 1989).

N.B. If you require a weighted linear regression then please use the multiple linear regression function in Stats
Direct; it will allow you to use just one predictor variable i.e. the simple linear regression situation. Note also that
the multiple regression option will also enable you to estimate a regression without an intercept i.e. forced through
the origin.

NVSU-FR-ICD-05-00 (081220) Page 29 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Example
From Armitage and Berry (1994, p. 161).
Test workbook (Regression worksheet: Birth Weight, % Increase).

The following data represent birth weights (oz) of babies and their percentage increase between 70 and 100 days
after birth.

Birth Weight % Increase


72 68
112 63
111 66
107 72
119 52
92 75
126 76
80 118
81 120
84 114
115 29
118 42
128 48
128 50
123 69
116 59
125 27
126 60
122 71
126 88
127 63
86 88
142 53
132 50
87 111
123 59
133 76
106 72
103 90
118 68
114 93
94 91

To analyze these data in Stats Direct you must first enter them into two columns in the workbook appropriately
labelled. Alternatively, open the test workbook using the file open function of the file menu. Then select Simple
Linear and Correlation from the Regression and Correlation section of the analysis menu. Select the column marked
"% Increase" when prompted for the response (Y) variable and then select "Birth weight" when prompted for the
predictor (x) variable.

For this example:


Simple linear regression

Equation: % Increase = -0.86433 Birth Weight +167.870079

Standard Error of slope = 0.175684


95% CI for population value of slope = -1.223125 to -0.505535

Correlation coefficient (r) = -0.668236 (r²= 0.446539)

NVSU-FR-ICD-05-00 (081220) Page 30 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

95% CI for r (Fisher's z transformed) = -0.824754 to -0.416618

t with 30 DF = -4.919791
Two sided P < .0001
Power (for 5% significance) = 99.01%

Correlation coefficient is significantly different from zero

From this analysis we have gained the equation for a straight line forced through our data i.e. % increase in weight
= 167.87 - 0.864 * birth weight. The r square value tells us that about 42% of the total variation about the Y mean
is explained by the regression line. The analysis of variance test for the regression, summarized by the ratio F,
shows that the regression itself was statistically highly significant. This is equivalent to a t test with the null
hypothesis that the slope is equal to zero. The confidence interval for the slope shows that with 95% confidence the
population value for the slope lies somewhere between -0.5 and -1.2. The correlation coefficient r was statistically
highly significantly different from zero. Its negative value indicates that there is an inverse relationship between X
and Y i.e. lower birth weight babies show greater % increases in weight at 70 to 100 days after birth. With 95%
confidence the population value for r lies somewhere between -0.4 and -0.8.

Graphs for Different Correlation Coefficients

Graphs always help bring concepts to life. The scatterplots below represent a spectrum of different correlation
coefficients. I’ve held the horizontal and vertical scales of the scatterplots constant to allow for valid comparisons
between them.

Correlation Coefficient = +1: A perfect positive relationship.

Correlation Coefficient = 0.8: A fairly strong positive relationship.

Correlation Coefficient = 0.6: A moderate positive relationship.

NVSU-FR-ICD-05-00 (081220) Page 31 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Correlation Coefficient = 0: No relationship. As one value increases, there is no tendency for the other
value to change in a specific direction.

Correlation Coefficient = -1: A perfect negative relationship.

Correlation Coefficient = -0.8: A fairly strong negative relationship.

NVSU-FR-ICD-05-00 (081220) Page 32 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Correlation Coefficient = -0.6: A moderate negative relationship.

V. LEARNING ACTIVITIES

1. Definition of Terms
a) Statistics
b) Data
c) Variable
d) Frequency
e) Population
f) Sample
g) Primary Data
h) Secondary Data
i) Raw Data
j) Discreet Data
k) Continuous Data
l) Probability Distribution
m) Space

Poor Fair Good


1 pts 3 pts 5 pts

The definition of the A definition of the The correct definition is


Vocabulary Definition vocabulary word is not vocabulary word is used, and is complete.
included on the page or incomplete.
the wrong definition is
written.

2. Identify different types of data.


3. Enumerate how to gather data.
4. Give different data and classify each. ( at least 10 )
5. Tally the given data by using the master sheet then rank the data
48 83 89 52 60 70 66 68 77 88 56 41 50 59
92 96 58 60 74 97 62 76 47 86 71 49 67 98
91 87 66 96 64 94 53 75 55 59 68 90 44 71
81 89 77 51 60 57 80 91

Solve for the a. mean, median, mode b. P50, P30, P70 c. D4, D8 d. Q2, Q3

6. Construct a frequency distribution table and solve for the informative columns included

NVSU-FR-ICD-05-00 (081220) Page 33 of 34


Republic of the Philippines
NUEVA VIZCAYA STATE UNIVERSITY
Bambang, Nueva Vizcaya
INSTRUCTIONAL MODULE
IM No.4:GEMath-1S-2020-2021

Data f Exact Mid <cf >cf <RCF >RCF RF %RF


UL LL limits point
71-80 1
61-70 1
51-60 9
41-50 13
31-40 10
21-30 16
11-20 5
n=55

VI. REFERENCES

A) Book/Printed Resources
Marie-Franie J. Frany et al. Fundamentals of Probability and Statistics for Engineering
Ronald Walpole, Ye et al. Probability and Statistics for Engineers and Scientist 9 th edition
Arciaga, Magcuyao. Statistics and Probability 1 st edition 2016
Reyes, Jocelyn L., et. al. (YEAR) ,Mathematics in the Modern World, Panday-Lahi Publishing
House, Inc., 2018
Nocon, Rizaldi C., et al. (YEAR) Essential Mathematics for the Modern World, C & E
Publishing, Inc.
2018
Punsalan, Twila G., et. al.(YEAR) Statistics A Simplified Approach, Rex Book Store
Paguio, Darwin P.,Statistics With Computer Based Discussions, Jimczyville Publications,
2012
Adam, John A. Mathematics in Nature: Modelling Patterns in the Natural World
Adam, John A. Mathematical Nature Walk
Aufman, R. et al. mathematical Excursions ( Chaps 1,2,3,4,5,8,11, and 13) 3 rd Ed (International
Edition)
COMAO Inc. For all Practical Purposes, Introduction to Contemporary Mathematics, 2 nd Ed.
Fisher, Carol Burns, The Language of Mathematics
Fisher, Carol Burns, The Language and Grammar of Mathematics
Hersh, R., What is Mathematics Really? (Chaps. 4 & 5)
Johnson and Mowry. Mathematics a Practical Odyssey ( Chap 12)
Moser and Chen. A Student Guide to Coding and Information Theory
Stewart, Ian. Nature’sw Numbers
Vistro-Yu, C. Geometry: Shapes, Patterns and Designs

B) e-Resources
[Link]

NVSU-FR-ICD-05-00 (081220) Page 34 of 34

You might also like