0% found this document useful (0 votes)
2 views17 pages

Statistical Analysis With Software Application (Stats)

The document provides an overview of statistical analysis, including the definitions of statistics, types of data, and the importance of statistics in decision-making across various fields such as business, healthcare, and marketing. It outlines key concepts such as descriptive and inferential statistics, sampling methods, and data presentation techniques using software like Excel. Additionally, it discusses the levels of measurement and the distinction between population and sample data.

Uploaded by

Leila Lorenzo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views17 pages

Statistical Analysis With Software Application (Stats)

The document provides an overview of statistical analysis, including the definitions of statistics, types of data, and the importance of statistics in decision-making across various fields such as business, healthcare, and marketing. It outlines key concepts such as descriptive and inferential statistics, sampling methods, and data presentation techniques using software like Excel. Additionally, it discusses the levels of measurement and the distinction between population and sample data.

Uploaded by

Leila Lorenzo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

STATISTICAL ANALYSIS WITH SOFTWARE APPLICATION

(STATS)

■​ Present data (ex. Tables and graphs)

MODULE 0: INTRODUCTION ■​ Characterize data (ex. Sample


mean)
WHAT IS STATISTICS?
●​ Statistics – the science of collecting,
organizing, analyzing, interpreting, and 37% lower poverty rate than last
presenting data. (1) collect, (2) organizing year, describing data lang
and presenting, (3) analyzing, connect ○​ Inferential Statistics
sample to real world (population) ■​ Generalizing from a sample to a
●​ Statistic – is a single measure, reported as a population, estimating unknown
number, used to summarize a sample data population parameters, drawing
set, for example,the average height of conclusions, and making decisions.
students in a university any sample is ■​ Drawing conclusions and/or
statistics; 3 per row (sample), get average making decisions concerning a
(statistic), percentage of a sample (statistic), population based only on sample
percentage of a population (parameter) data may conclusions talaga,
●​ The branch of mathematics that merong binubuong opinion
transforms data into useful information for ■​ Estimation
decision makers ●​ Ex. estimate the population
●​ Examples mean weight using the sample
○​ Average height for the length of the mean weight
gowns ■​ Hypothesis testing
○​ Maximum height to design the height ●​ Ex. test the claim that the
of the doorways of the classrooms, etc. population mean weight is 120
●​ Kinds of Statistics pounds.
○​ Descriptive Statistics ■​ Drawing conclusions and/or
■​ The collection, presentation, and making decisions concerning a
summary of data (either using population based on sample
charts and graphs or using results.
numerical summary). TASKS IN STATISTICS
■​ Collecting, summarizing, and
describing data
■​ Collect data (ex. Survey)
○​ Process improvement – large firms
have formal systems for continuous
quality improvement. Statistics helps
firms oversee their suppliers, monitor
their internal operations, and identify
problems. Quality improvement goes
far beyond statistics, but every college
WHY STUDY STATISTICS?
graduate is expected to know enough
●​ Decision makers use statistics to:
statistics to understand its role in
○​ Present and describe business data and
quality improvement.
information properly
○​ Auditing – the firm has learned that
○​ Draw conclusions about large
some invoices are being paid
populations, using information
incorrectly, but it doesn't know how
collected from samples
widespread the problem is. A sample of
○​ Make reliable forecasts about a business
invoices can be used to estimate the
activity
proportion of incorrectly paid invoices.
○​ Improve business processes
○​ Marketing – many companies use
●​ Reasons to study statistics
Customer Relationship Management
○​ Communication – understanding the
(CRM) to analyze customer data from
language of statistics facilitates
multiple sources. With statistical and
communication and improves problem
analytics tools such as correlation and
solving
date mining, they identify specific
○​ Computer skills – the use of
needs of different customer groups,
spreadsheets for data analysis and word
and this helps them market their
processors or presentation software for
products and services more effectively.
reports improves upon your existing
○​ Health care – evaluate 100 incoming
skills
patients using a 42-item physical and
○​ Information management – statistics
mental assessment questionnaire.
help summarize small and large
○​ Quality improvement – initiate a
amounts of information (data) and
triple inspection program, setting
reveal underlying relationships
penalties for workers who produce
○​ Technical literacy – career
poor-quality output
opportunities are in growth industries
○​ Purchasing – a food producer
propelled by advanced technology. The
purchases plastic containers for
use of statistical software increases your
packaging its product. Inspection of
technical literacy
the most recent shipment of 500
containers found that 3 of the ●​ Population – measures used to describe
containers were defective. The the population
supplier's historical defect rate is .005. ●​ Sample – measures computed from sample
Has the defect rate really risen or is this data
simply a “bad” batch?
○​ Medicine – determine whether a new
drug is really better than the placebo or
if the difference is due to chance.
○​ Operations management – manage
inventory by forecasting consumer
demand.
○​ Product warranty – determine the
average dollar cost of engine warranty
claims on a new hybrid engine.
BASIC VOCABULARY OF STATISTICS
●​ Variable – a characteristic of an item or
individual.
●​ Data – the different values associated with
a variable
●​ Operational definitions – variable values
are meaningless unless their variables have
operational definitions, universally
accepted meanings that are clear to all
associated with an analysis. nondictionary
definition
●​ Population – a population consists of all
the items or individuals about which you
want to draw a conclusion.
●​ Sample – the portion of a population
selected for analysis
●​ Parameter – a numerical measure that
describes a characteristic of a population
●​ Statistic – a numerical measure that
describes a characteristic of a sample.
POPULATION VS SAMPLE
TYPES OF VARIABLES

MODULE 1A: REVIEW OF BASIC ●​ Categorical (qualitative) variables –


STATISTICAL CONCEPTS (DATA have values that can only be placed into
MEASUREMENT) categories, such as “yes” and “no.” textual;
TIN, SSS; cant perform operations
WHY COLLECT DATA? ●​ Numerical (quantitative) variables –
●​ A marketing research analyst needs to have values that represent quantities. can
assess the effectiveness of a new television perform operations; there are numerical in
advertisement. form but shd be performed in categorical;
●​ A pharmaceutical manufacturer needs to age,
determine whether a new drug is more
effective than those currently in use.
●​ An operations manager wants to monitor a
manufacturing process to find out whether
the quality of product being manufactured
is conforming to company standards.
●​ An auditor wants to review the financial
transactions of a company in order to
determine whether the company is in
discrete – no of passenger in an airplane, no
compliance with generally accepted
of female; can count them,
accounting principles.
Continuous – height, weight, income, ROA,
SOURCES OF DATA
ROI; can measure them, with instrument,
●​ Primary sources ikaw mismo; firsthand
formula to get
○​ The data collector is the one using the
●​ The difference between numerical and
data for analysis
categorical data
■​ Data from a market survey
■​ Data collected from an experiment
■​ Observed data
●​ Secondary sources birth cert
○​ The person performing data analysis is
not the data collector
■​ Analyzing census data summary
■​ Examining data from print ○​ Note: Ambiguity is introduced when
journals, from government continuous data are rounded to whole
agencies or data published on the numbers. Be cautious
internet.
●​ The levels of measurement in data and ○​ Interval Scale; Interval
ways of coding data scales of Measurement ranking, no true value of
measurement; ratio (most powerful), (2) 0, cant multiply
Interval), (3) Ordinal, (4) Nominal; if you ■​ An ordered scale in which the
have higher data, you can go down (like difference between measurements
ratio to ordinal pwede, bawal pataas) is a meaningful quantity but the
measurements do not have a true
zero point
■​ Data can only be ranked, but also
have meaningful intervals between
scale points (e.g., difference
between 60°F and 70°F is same as
LEVELS OF MEASUREMENT difference between 20°F and 30°F)
●​ Categorical variables ■​ Since intervals between numbers
○​ Nominal scale name ang root word, represent distances, mathematical
pang identify lang, cant perform operations can be performed (e.g.,
operations; ex. sole prop, corp, part, coop average)
■​ Classifies data into distinct ■​ Zero point of interval scales is
categories in which no ranking is arbitrary, so ratios are not
implied meaningful (e.g., 60°F is not twice
as warm as 30°F)
■​ Likert scales
●​ A special case of interval data
○​ Ordinal scale may rank, position, w/o frequently used in survey
meaningful difference research.
■​ Classifies data into distinct ●​ The coarseness of a Likert scale
categories in which ranking is refers to the number of scale
implied points (typically 5 or 7) if no
value/number, cinoconsider sa
iba na ordinal ang likert scale

●​ Numerical variables
●​ Neutral midpoint (Neither
agree nor disagree) – allowed
if an odd number of scale quantity and the measurements
points is used or omitted to have a true zero point.
force the respondent to “lean” ■​ Ratio data – have all properties of
one way or the other. nominal, ordinal, and interval data
●​ Likert data are coded types and also possess a meaningful
numerically (e.g., 1-5) but any zero (absence of quantity being
equally spaced values will work measured).
■​ Because this zero point, ratios of
data values are meaningful (e.g.,
$20m profit is twice as much as
$10m)
■​ Zero does not have to be observable
in the date, it is an absolute
■​ Careful choice of verbal anchors reference point.
results in measurable intervals
(e.g., the distance from 1 to 2 is
“the same” as the interval from 3 to
4)
■​ Ratios are not meaningful (here 4
is not twice 2)
●​ Use the following procedure to
■​ Many statistical calculations can be
recognize data types:
performed (e.g., averages,
○​ Is there a meaningful zero point?
correlations, etc)
■​ Ratio data – all statistical
■​ More variants
operations are allowed
○​ Are intervals between scale points
meaningful?
■​ Interval data – common statistics
allowed, e.g., means and standard
○​ Ratio Scale; Ratio Measurement deviations
annual income (pwede 0), can do ○​ Do scale points represent rankings?
operations ■​ Ordinal data – restricted to
■​ An ordered scale in which the certain types of nonparametric
difference between the statistical tests
measurements is a meaningful ○​ Are there discrete categories?
■​ Nominal data – only counting ○​ Sample – involves looking only at
allowed, e.g., finding the mode some items selected from the
●​ Changing data by recording population
○​ In order to simplify data or when exact ○​ Census – an examination of all items
data magnitude is of little interest, ratio in a defined population census is to
data can be recorded downward into parameter, sample is to statistic
ordinal or nominal measurements (but ○​ Why can the United States Census
not conversely) survey every person in the
■​ For example, recode systolic blood population? – mobility,
pressure as “normal” (under 130), undocumented workers, budget
“elevated” (130 to 140), or “high” constraints, incomplete responses, etc.
(over 140).
■​ The above recoded data are ordinal
(ranking is preserved), but intervals
are unequal and some information
is lost.
SAMPLING CONCEPTS
●​ Sample or population?
○​ Population – involves all of the items G power analysis –
one is interested in. It may be finite ●​ Parameter or statistic?
(e.g., all of the passengers on a plane) or ○​ Parameter – a measurement or
effectively infinite (e.g., all of the Cokes characteristic of the population
produced in an ongoing bottling ■​ Usually unknown because we can
process). rarely observe the entire population
○​ Sample – a subset of the population ○​ Statistic – a numerical value
and involves looking only at some of computed from a sample
the items selected from the population ■​ Can be used as estimates of
parameters found in the
population
○​ Symbols – used to represent
population parameters and sample
statistics.
○​ From a sample of n items, chosen from
●​ Sample or census?
a population, we compute statistics
that can be used as estimates of
parameters found in the population
○​ To avoid confusion, we use different
symbols for each parameter and its
corresponding statistic.
○​ Thus, the population mean is denoted
by μ (the lowercase Greek letter mu)
while the same mean is x̄ .
○​ The population proportion is denoted
π (the lowercase Greek letter pi), while
the sample proportion is p.
●​ Target population
○​ The population must be carefully
specified and the sample must be
drawn scientifically so that the sample
is representative.
○​ Target population – the population
we are interested in (e.g., US gasoline
prices).
○​ Sampling frame – the group from
which we take the sample (e.g., 115,000
stations).
○​ The frame should not differ from the
target population
MODULE 1B: DATA PRESENTATION
USING EXCEL
●​ Bar chart – a bar shows each category, the
BENEFITS OF EXCEL’S ATTRACTIVE
length of which represents the amount,
FEATURES
frequency or percentage of values falling
●​ Not having to incur the extra costs of using
into a category. Length of bar, largest –
specialized statistical programs
more frequent
●​ Familiarity with excel
●​ Easy to use and easy to learn
●​ Allows to use the same worksheet-based
data that users have created for other
business purposes
●​ Some graphical functions produce more
vivid visual outputs.
●​ Pie chart – a circle broken up into slices
UNIVARIATE AND BIVARIATE DATA
that represent categories. The size of each
●​ Univariate data – set of n observations in
slice of the pie varies according to the
one variable, e.g., height of Grade 3 pupils,
percentage in each category only use
GWA of second year students, etc.
percentage
●​ Bivariate data – a set of n observations
involving 2 variables, g.g., height and
weight of newly born babies, monthly
income and expenses of selected employees.
Multivariate data – 3 or more variables
ORGANIZING UNIVARIATE
ORGANIZING BIVARIATE
CATEGORICAL DATA
CATEGORICAL DATA
●​ Summary table – indicates the frequency,
●​ Cross Tabulations
amount, or percentage of items in a set of
○​ Cross-classification (contingency)
categories so that you can see differences
table
between categories. Apa style – no vertical
■​ Presents the results of 2 categorical
lines, puro horizontal lang; 50 sample size;
variables. The joint responses are
in reporting, highest and lowest lang, engage
classified so that the categories of
the audience; can be converted to bar chart
one variable are located in the rows
and the categories of the other
variable are located in the columns.
■​ Cell – is the intersection of the to the right on each row. Head, lead – stem;
row and column and the value in tail – leaves; teens (1 ipair mo sa 6, 16 17
the cell represents the data 18); can only do manually, no excel
corresponding to that specific function; for small sample sizes; if decimals,
pairing of row and column have to round them off
categories.

○​ Side-by-side bar charts


■​ A useful way to visually display the ●​ Frequency distribution and histogram
results of cross-classification data is ○​ Bins and bin limits
by constructing a side-by-side bar ■​ Frequency distribution – a table
chart. formed by classifying n data values
into k classes (bins) k – groupings,
class groupings, class intervals; n –
class size, bin limits,
■​ Bin limits – define the values to be
included in each bin. Widths must
ORGANIZING NUMERICAL DATA all be the same
(UNIVARIATE DATA) ■​ Frequencies – the number of
●​ Ordered array – a sequence of data, in observations within each bin
rank order, from the smallest value to the vertical distance/difference against
largest values manageable/small amt of one another, shd be constant
data if around 50 lang ganon; if too ■​ Express as relative frequencies
many(e.g., thousands), not advisable (frequency divided by the total) or
percentages (relative frequency
times 100)
○​ Constructing a frequency
distribution 1st: now n

●​ Stem and leaf display – organizes data


into groups (called stems) so that the values
within each group (the leaves) branch out
■​ Find the smallest and largest data ■​ Put the data values in the
values appropriate bin
■​ Create the table, you can include:
●​ Frequencies – counts for each
bin
●​ Relative frequencies –
■​ Choose the number of bins (k) no
absolute frequency divided by
of groupings
the total number of data
●​ K should be much smaller than
●​ Cumulative frequency –
n cuz n is the total
accumulated relative frequency
●​ Too many bins results in
values as bin limit increases.
sparsely populated bins, too
●​ The histogram (bar chart) visual display;
few and dissimilar data values
appears to be a bar chart pero diff; a kind of
are lumped together.
a bar chart
●​ Herbert Sturges proposes the
○​ A graph of the data in a frequency
following rule:
distribution
○​ Class boundaries (or class
midpoints) – shown on the horizontal
axis
○​ Vertical axis – either frequency,
​ His formula (if more than
relative frequency, or percentage
1024):
○​ Bars of the appropriate heights are used
​ 1 + 3.322(log n) then round up
to represent the number of
■​ Set the bin limits bin width = class
observations within each class.
size; highest – lowest data, divide by
no of grpings

For example, for k=7 bins, the


20 21 22 23 24 25 = 6 data within 1
approximate bin width is:
interval
○​ In Excel
■​ Select Tools/Data Analysis
6 numerical data within class
interval
represent the data in that class and then
connecting the sequence of midpoints
at their respective class percentages
she’ll also use class midpoint as X and
frequency sa Y
○​ Cumulative percentage polygon or
ogive – displays the variable of interest
along the X axis, and the cumulative
■​ Choose histogram percentages along the Y axis Y dapat
cumu %
○​ Note: in a percentage polygon the
vertical axis would be defined to show
the percentage of observations per class
Some cases, class midpoint = class mark,
■​ Input data range and bin range average of upper and lower limit
(bin range is a cell range containing
the upper class boundaries for each
class grouping <20 use 19.9) diff
ways to input range, dito, upper class
boundaries (midway between the
upper and lower limit of the
ORGANIZING NUMERICAL DATA
succeeding interval presented in
(BIVARIATE DATA)
descending order; if given 20 as
●​ Scatter plots
upper limit, can use 19.9) ang
○​ Used for numerical data consisting of
gamit
paired observations taken from 2
■​ Select chart output and click “ok”
numerical variables
○​ 1 variable is measured on the vertical
axis and the other variable is measured
on the horizontal axis x axis –
independent variable, left side of the
column; y axis – dependent variable,
right side of the column; for
correlational analysis
●​ Polygon (Line Graph)
○​ Percentage polygon – formed by
having the midpoint of each class
○​ In excel (1997-2003) ○​
■​ Select the chart wizard
■​ Select XY(Scatter) option, then
click “next”
■​ When prompted, enter the data
range, then click a’next”
■​ Enter title, axis labels, and legend, ○​
and click “finish” PRINCIPLES OF EXCELLENT GRAPHS
●​ The graph should not distort the data
maganda simple, not yung 3d na d naman
maganda
●​ The graph should not contain unnecessary
adornments (sometimes referred to as
chart junk)
●​ The scale on the vertical axis should begin
at zero
■​
●​ All axes should be properly labeled
●​ Time series plot
●​ The graph should contain a title
○​ Used to study patterns in the values of
●​ The simplest possible graph should be used
a numerical variable over time. Each
for a given set of data
value is plotted as a point in 2
GRAPHICAL ERRORS
dimension with the time period on the
●​ Chart junk
horizontal X axis and the variable of
interest on the Y axis

●​ No relative basis
●​ No zero point in the vertical axis

CHART SUMMARY

Pareto – focus on important categories, separates


vital few from the trivial many, few many are
spread out over a large number of categories,
application can be found in factory when
dealing w/ data of nonconforming items
■​ Excel formula

MODULE 1C: MEASURES OF =AVERAGE(Data)


CENTRAL TENDENCY, POSITION, ■​ Pro – familiar and uses all the
AND, VARIABILITY sample information
■​ Con – influenced by extreme
SUMMARY DEFINITIONS values
●​ Central tendency – the extent to which ○​ Median
all the data values group around a typical or ■​ Middle value in sorted array
central value ■​ Excel formula =MEDIAN(Data)
●​ Variation – the amount of dispersion, or ■​ Pro – robust when extreme data
scattering, of values values exist
●​ Shape – the pattern of the distribution of ■​ Con – ignores extremes and can be
values from the lowest value to the highest affected by gaps in data values.
value. ○​ Mode
NUMERICAL DESCRIPTION ■​ Most frequently occurring data
●​ 3 characteristics of numerical data (and value
its interpretation) ■​ Excel formula =MODE(Data)
○​ Central tendency ■​ Pro – useful for attribute data or
■​ Where are the data values discrete data with a small range
concentrated? What seem to be ■​ Con – may not be unique, and is
typical or middle data values? not helpful for continuous data
○​ Dispersion or variation ○​ Trimmed mean
■​ How much variation is there in the ■​ Same as the mean except omit
data/ how spread out are the data highest and lowest k% of data
values? Are there unusual values? values (e.g., 5%)
○​ Shape ■​ Excel formula
■​ Are the data values distributed =TRIMMEAN(Data, Percent)
symmetrically? Skewed? Sharply ■​ Pro – mitigates effects of extreme
peaked? Flat? bimodal? values
CENTRAL TENDENCY ■​ Con – excludes some data values
●​ 5 measures of central tendency that could be relevant.
○​ Mean ○​ Geometric mean
■​ Formula ■​ Formula
■​ Excel formula ○​ Example
=GEOMEAN(Data)
■​ Pro – useful for growth rates and
mitigates high extremes
■​ Con – less familiar and requires
positive data.
MEASURES OF CENTRAL TENDENCY
THE ARITHMETIC MEAN
●​ The most common measure of central
tendency.
●​ For a sample of size n:

●​ The most common measure of central


GROWTH RATES
tendency
●​ The average growth rate is given by taking
●​ Mean – sum of values divided by the
the geometric mean of the ratios of each
number of values
year’s revenue to the preceding year
●​ Affected by extreme values (outliers)
●​ Due to cancellations, only the first and last
year are relevant.

THE GEOMETRIC MEAN


THE MEDIAN
●​ Geometric mean
●​ In an ordered array, the median is the
○​ Used to measure the rate of change of a
“middle” number (50% above, 50% below)
variable over time
●​ Not affected by extreme values

●​ Geometric mean rate of return


○​ Measures the status of an investment ●​ Locating the median
over time ○​ The median of an ordered set of data is
○​ Where Ri is the rate of return in time located at the (n+1)/2 ranked value
period i
○​ If the number of values is odd, the
median is the middle number
○​ If the number of values is even, the
median is the average of the 2 middle
numbers
○​ NOTE: the (n+1)/2 is NOT the value ●​
of the median, only the position of the ●​
median in the ranked data. ●​
THE MODE ●​
●​ Value that occurs most often ●​
●​ Not affected by extreme values ●​
●​ Used for categorical data
●​ Used for numerical primarily when
grouped
●​ There may be no mode
●​ There may be several modes

TRIMMED MEAN
●​ To calculate the trimmed mean, first
remove the highest and lowest k percent of
the observations
●​ To determine how many observations to
trim, multiply k by n and round of the
result
●​ Let us say that k*n=3.4=3. So, we would
remove the 3 smallest and 3 largest
observations before averaging the
remaining values
EXAMPLE

You might also like