0% found this document useful (0 votes)
45 views4 pages

Engineering Data Analysis Overview

This document discusses key concepts in engineering data analysis including: 1. It defines data, variables, and different types of variables such as dependent and independent. 2. It outlines different types of data including qualitative (categorical) data and quantitative data, and levels of data measurement from nominal to ratio. 3. It discusses populations and samples, describing finite vs infinite populations and different sampling methods. 4. It covers frequency distributions including ungrouped, grouped, and relative frequencies, and how to make a basic frequency table to organize data.

Uploaded by

Alice Co
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
45 views4 pages

Engineering Data Analysis Overview

This document discusses key concepts in engineering data analysis including: 1. It defines data, variables, and different types of variables such as dependent and independent. 2. It outlines different types of data including qualitative (categorical) data and quantitative data, and levels of data measurement from nominal to ratio. 3. It discusses populations and samples, describing finite vs infinite populations and different sampling methods. 4. It covers frequency distributions including ungrouped, grouped, and relative frequencies, and how to make a basic frequency table to organize data.

Uploaded by

Alice Co
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

ENGINEERING DATA ANALYSIS

DATA – a collection of discrete values that convey


information, describing, quantity, facts or statics.
STATICS – is the study of analysis, presentation,
collection, interpretation, organization and large 2 TYPES OF DATA
data presentation. It can be defined as a function 1. QUALITATIVE (categorical)
of the given data. • Describes the object under consideration
using a finite set of discrete classes.
VARIABLE – a mathematical symbol that may
• Can’t be counted or measured easily
represent a number, a vector, a matrix, a function,
using numbers.
the argument of the function, a set of an element of
Ex. gender of a person
the set
a. NOMINAL – set of values that don’t
– Latin word “variabilis” – changeable
possess a natural ordering
OTHER SPECIFIC NAME FOR VARIABLES: Ex. hair color
1. UNKNOWN – a variable in an equation which b. ORDINAL – types of values that have a
has to be solved for. natural ordering while maintaining their class
2. INDETERMINATE – is a symbol commonly of values.
called variable that appears in a polynomial or Ex. size of a clothing brand
a formal power of series. 2. QUANTITATIVE
a. POLYNOMIAL – an expression consisting • Tries to quantity things and does by
of indeterminates and coefficients that considering numerical values.
involves only the operation of addition, a. DISCRETE
subtraction, multiplication, and positive
• consist of numerical variables that are
integers power of variables.
easily counted.
Ex. X – 4x + 7
2
• often identified through graphs.
variable of polynomial Ex. integers or whole numbers
b. PARAMETER – a quantity (usually
numbers) which is a part of the input of a b. CONTINUOUS
problem, and remains constant during the
• has numerical variable with an infinite
whole solution of the problem.
number of collected values.
Ex. F(x) = ax2 + bx + c Variable
of variable of a function Ex. height, temperature, weight
parameter/constant
LEVEL OF DATE MEASUREMENT:
DEPENDENT AND INDEPENDENT VARIABLE 4 LEVELS:
1. DEPENDENT (OUTPUT) – a variable that is • NOMINAL – data can be categorized
implicitly a function of another variable. Ex. city of birth, ethnicity
2. INDEPENDENT (INPUT) – a variable (often • ORDINAL – data can be categorized and
denoted by x) whole variation does not on that rank
of another. Ex. top 5 Olympic medalist
Ex. time, space, mass, density • INTERVAL – data can be categorized, rank
and evenly spaced.
ENGINEERING DATA ANALYSIS
Ex. test scores, temperature • describes the number of observations for
• RATIO – data can be categorized, rank, each possible value of a variable.
evenly spaced and has a natural zero. • depicted using graphs and frequency
tables.
Ex. height, weight, age
FREQUENCY OF A VALUE
POPULATION & SAMPLES
• the number of times it occurs in a dataset.
POPULATION – includes all the elements from the
data set and measurable characteristics of the TYPES OF FREQUENCY DISTRIBUTION:
population such as mean and standard deviation. 1. UNGROUPED – the number of observations of
each value of a variable
TYPES OF POPULATION:
– used for categorical
1. FINITE – also known as countable population
variables.
in which population can be counted.
2. GROUPED – the number of observations of
– the population of all individuals or
each class interval of a variable.
objects that are finite.
• CLASS INTERVAL – are ordered
Ex. employees of a company
groupings of a variable's values.
2. INFINITE – also known as uncountable
3. RELATIVE – the proportion of observation of
population in which the counting of units in the
each value or class interval of a variable.
population is not possible
– used for any type of variable especially when
Ex. numbers of germs
comparing frequencies.
3. EXISTENT – population of concrete individuals
4. CUMULATIVE – the sum of frequencies less
Ex. books, students
than or equal to each value or class interval of
4. HYPOTHETICAL – whose unit is not available
a variable.
in solid form
– used for ordinary
Ex. outcome of rolling a dice
or quantitative variables when understanding how
SAMPLES – includes one or more observations often observations fall below certain values.
that are drawn from the population
➢ HOW TO MAKE A FREQUENCY TABLE?
• SAMPLING – the process of selecting
FREQUENCY TABLE – an effective way to
the sample from the population
summarize or organize a dataset.
1. PROBABILITY – the population units
– usually composed of two
cannot be selected.
columns
Ex. simple random, stratified, cluster,
– values or class interval
disproportionate
– their frequencies
2. NON-PROBABILITY – the population
a. UNGROUPED
units can be selected
1. create a table
Ex. quota, purposive, judgement
VARIABLE’S NAME FREQUENCY
FREQUENCY DISTRIBUTION
• are visual display that organize and present
frequency counts so that the information can be
interpreted more easily.
ENGINEERING DATA ANALYSIS
• ORDINAL VARIABLES – the values GRAPH OF A FREQUENCY DISTRIBUTION
should be ordered from smallest to 1. PIE CHART – a circular graph that shows the
largest value. relative frequency distribution of a nominal
• NOMINAL VARIABLES – the values variable.
can be in any order in the table. 2. BAR CHART – a graph that shows frequency
2. count the frequencies or relative frequency distribution of a
• the frequencies are the number of categorical variable (nominal or ordinal)
times each value occur – y-axis - frequencies / relative
• if dataset is large, count frequencies frequencies
by tallying. – x-axis - values
b. GROUPED 3. HISTOGRAM – a graph that shows the
1. divide the variable into class intervals frequency or relative frequency distribution of a
a. calculates the range: subtract the quantitative variable
lowest value in the dataset from the – y-axis - frequencies / relative
highest. frequencies
b. decides the class interval width – x-axis - interval class
c. calculates the class intervals
2. create a table BAR CHART HISTOGRAPH
3. count the frequencies TYPE OF Categorical Quantitative
c. RELATIVE VARIABLE
1. create an ungrouped or grouped VALUE Ungrouped Grouped
frequency table. GROUPING (values) (interval class)
2. add a third column to the table for BAR Can be a Never a space
relative frequencies SPACING space between bars
o to calculate the relative frequency, between bars
divide each frequency by the sample BAR ORDER Can be in any Can only be
size. the sample size is the sum of order ordered from
frequencies. lowest to
d. CUMULATIVE highest
1. create an ungrouped or grouped MEAN, MEDIADE & RANGE
frequency table.
2. add a third column to the table for • MEAN – the total number of all
cumulative frequency values divided by the number of the values.
o cumulative frequency is the number
of observations less than or equal to or

a certain value or class interval. • MEDIAN – is the middle number in

3. optional: cumulative relative frequency a list of numbers ordered lowest to highest.

o divide each cumulative frequency by u=


the sample size. • MODE – is the value appear most
often in a set of data.
ENGINEERING DATA ANALYSIS

Mo )h
Where: L – lower limit
H – size of the class interval
F1 – first frequency
F0 – preceding frequency
F2 – succeeding frequency
• RANGE – is the difference between
the lowest and highest value.
Range = max. value – min. value
• VARIANCE – is the expectation of
the squared deviation of a random variable
from its population mean or sample mean.

or
Where: 𝒔𝟐/𝝈𝟐 – sample / population variable
𝒙𝟏 – the value of the one
observation
𝐱̅ /𝑴 – mean value
𝒏 − 𝟏 / 𝑵 – number of observations
• STANDARD DEVIATION – is the
measure of the distribution of the statistical
data.
Σ (𝑥 1 − 𝑀 )2
or 𝜎=√ 𝑁

Common questions

Powered by AI

A frequency table condenses data by summarizing the occurrences of each value or range, reducing data size, and presenting it in a tabular form that highlights the distribution and patterns such as central tendency, dispersion, and anomalies. By allowing easier comparisons and visualization through charts or graphs, it facilitates intuitive understanding of complex datasets.

A dependent variable is one that depends on and is influenced by an independent variable. It is essentially the output of a function, determined by changes in the independent variable, which is the input that varies without being affected by other variables within the function.

Central tendency offers statistical metrics to identify the center of a dataset. The mean provides an average, showing the balance point of the data. The median gives the middle value, useful in skewed distributions to denote central location. The mode indicates the most frequently occurring value, highlighting peaks in data density. Together, they offer comprehensive insights into the data center, frequency, and distribution shape.

Grouped frequency tables are more advantageous when the dataset consists of a large range of values or continuous variables because they simplify and enhance the visualization by organizing data into class intervals. This reduces complexity and allows for clearer trends or patterns, unlike ungrouped tables that may overwhelm with individual values when dealing with large datasets.

Levels of data measurement—nominal, ordinal, interval, and ratio—determine the permissible types of statistical operations that can be conducted. Nominal data allows only for categorization, ordinal adds ranking, interval introduces equal spacing allowing for meaningful subtraction, and ratio includes a true zero point enabling all arithmetic operations. This progression of measurement levels facilitates increasingly complex analyses, from simple frequency counts to sophisticated modeling.

Standard deviation is often preferred over variance because it is in the same units as the data, making it more interpretable especially when comparing with mean values. Variance, being in squared units, can exaggerate the perception of spread, whereas standard deviation provides an accessible measure directly comparable to individual data points and the mean.

Probability sampling methods ensure each unit has an equal chance of selection, which reduces selection bias and enhances representativeness, allowing for generalization of results to the entire population. Non-probability sampling may lead to biased selections and is less suitable for generalization, but it is often easier and cheaper to implement when population lists are unavailable.

Finite populations allow for complete enumeration or direct application of sampling without omission, facilitating exact metrics like total size and direct variance computation. Infinite populations necessitate sampling with assumptions about distribution, typically using central limit theorem principles, to make probabilistic inferences due to impracticality of exhaustive analysis. This distinction influences sampling design, margin of error calculation, and result generalizability.

A cumulative frequency distribution is helpful in understanding the accumulation of data points up to a certain value, which is essential in determining how often observations fall below a certain threshold. This is particularly useful in analyses where the relative standing or ranking within a dataset is more important than the proportion of each single value, such as in percentile calculations.

Nominal variables do not have a natural order, exemplified by characteristics like hair color or city of birth. In contrast, ordinal variables have a natural order, such as clothing sizes or rankings, which can be ordered but not evenly spaced.

You might also like