STATISTICAL ANALYSIS WITH SOFTWARE
APPLICATION
References:
PowerPoint Presentation
Supplementary:
Basic Business Statistics. Chapter 1. Defining and Collecting Data
Basic Business Statistics. Chapter 2. Organizing and Visualizing Variables
CHAPTER 1: DEFINING AND COLLECTING DATA
Introduction to Business Statistics
Operational Definitions of Terms
Variable A characteristic of an item or individual.
Data The set of individual values associated with a variable.
Statistics The methods that help transform data into useful information for decision makers.
TYPES OF VARIABLES (DCOVA)
a. Categorical (Qualitative)
- variables have values that can only be placed into categories, such as “agree” and “disagree.”
b. Numerical (Quantitative)
- variables have values that represent a counted or measured quantity.
● Discrete variables arise from a counting process
● Continuous variables arise from a measuring process
LEVELS OF MEASUREMENT
a. Nominal Scale (Categorical)
- Classifies data into distinct categories in which no ranking is implied.
b. Ordinal Scale (Categorical)
- Classifies data into distinct categories in which (with) ranking is implied.
c. Interval Scale (Numerical)
- is an ordered scale in which the difference between measurements is a meaningful quantity
but the measurements do not have a true zero point.
d. Ration Scale (Numerical)
- is an ordered scale in which the difference between the measurements is a meaningful
quantity and the measurements have a true zero point.
SOURCES OF DATA (DCOVA)
a. Primary Sources
- The data collector is the one using the data for analysis
● Data from a political survey
● Data collected from an experiment
● Observed data
b. Secondary Sources
- The person performing data analysis is not the data collector
● Analyzing census data
● Examining data from print journals or data published on the internet
Data is collected from either a…
a. Population
- consists of all the items or individuals about which you want to draw a conclusion. The
population is the “large group”
b. Sample
- is the portion of a population selected for analysis. The sample is the “small group
TYPES OF SAMPLES
Convenience
Nonprobability Sample
Judgment
Simple Random Sample
- Simple to use; May not be a good
representation of the population’s
underlying characteristics
Systematic Sample (same with Simple)
Probability Sample Stratified Random Sample
- Ensures representation of individuals
across the entire population
Cluster Sample
- More cost effective; Less efficient (need
larger sample to acquire the same level
- of precision)
a. Nonprobability Sample
- In a nonprobability sample, items included are chosen without regard to their probability of
occurrence.
● Convenience Sampling
- items are selected based only on the fact that they are easy, inexpensive, or
convenient to sample
● Judgment Sampling
- you get the opinions of pre-selected experts in the subject matter.
b. Probability Sample
● Simple Random Sample
- Every individual or item from the frame
has an equal chance of being selected.
- Selection may be with replacement
(selected individual is returned to frame
for possible reselection) or without
replacement (selected individual isn’t
returned to the frame).
- Samples obtained from table of random
numbers or computer random number generators.
● Systematic Sample
- Decide on sample size: n
- Divide frame of N individuals into
groups of k individuals: k = N/n
- Randomly select one individual from the
1st group
- Select every kth individual thereafter
● Stratified Sample
- Divide population into two or more subgroups (called strata) according to some common
characteristic
- A simple random sample is selected from
each subgroup, with sample sizes
proportional to strata sizes
- Samples from subgroups are combined
into one
- This is a common technique when
sampling the population of voters, stratifying across racial or socio-economic lines.
● Cluster Sample
- Population is divided into several “clusters,” each representative of the population
- A simple random sample of clusters is selected
- All items in the selected clusters can be used, or items can be chosen from a cluster using
another probability sampling technique
- A common application of cluster sampling involves election exit polls, where certain election
districts are selected and sampled.
CHAPTER 2: ORGANIZING AND VISUALIZING VARIABLES
TABLES USED FOR ORGANIZING CATEGORICAL DATA
Summarizing Table
- Tallies the frequencies or percentages of items in a set of categories so that you can see
differences between categories.
Example:
Contingency Table
- Used to study patterns that may exist between the responses of two or more categorical
variables.
- Cross tabulates or tallies jointly the responses of the categorical variables.
- For two variables the tallies for one variable are located in the rows and the tallies for the
second variable are located in the columns.
Example:
● A random sample of 400 invoices is drawn.
● Each invoice is categorized as a small, medium, or large
amount.
● Each invoice is also examined to identify if there are any errors.
● This data is then organized in the contingency table to the right.
TABLES USED FOR ORGANIZING NUMERICAL DATA
a. Ordered Array
- Is a sequence of data, in rank order, from the
smallest value to the largest value.
- Shows range (minimum value to maximum value)
- May help identify outliers (unusual observations)
b. Frequency Distribution
- The frequency distribution is a summary table in which the data are arranged into numerically
ordered classes.
- You must give attention to selecting the appropriate number of class groupings for the table,
determining a suitable width of a class grouping, and establishing the boundaries of each class
grouping to avoid overlapping.
- The number of classes depends on the number of values in the data. With a larger number of
values, typically there are more classes. In general, a frequency distribution should have at
least 5 but no more than 15 classes.
- To determine the width of a class interval, you divide the range (Highest value–Lowest value)
of the data by the number of class groupings desired.
Why use a Frequency Distribution?
- It condenses the raw data into a more useful form
- It allows for a quick visual interpretation of the data
- It enables the determination of the major characteristics of the data set including
where the data are concentrated / clustered
Some Tips
- Different class boundaries may provide different pictures for the same data
(especially for smaller data sets)
- Shifts in data concentration may show up when different class boundaries are chosen
- As the size of the data set increases, the impact of alterations in the selection of class
boundaries is greatly reduced
- When comparing two or more groups with different sample sizes, you must use either
a relative frequency or a percentage distribution
Frequency Distribution Examples:
a. A manufacturer of insulation randomly selects 20 winter days and records the daily
high temperature:
24, 35, 17, 21, 24, 37, 26, 46, 58, 30, 32, 13, 12, 38, 41, 43, 44, 27, 53, 27
b. Sort raw data in ascending order:
12, 13, 17, 21, 24, 24, 26, 27, 27, 30, 32, 35, 37, 38, 41, 43, 44, 46, 53, 58
Find range: 58 - 12 = 46 Select number of classes: 5 (usually between 5 and 15)
Compute class interval (width): 10 (46/5 then round up)
Determine class boundaries (limits):
Class 1: 10 to less than 20
Class 2: 20 to less than 30
Class 3: 30 to less than 40
Class 4: 40 to less than 50
Class 5: 50 to less than 60
Compute class midpoints: 15, 25, 35, 45, 55
Count observations & assign to classes
c. Cumulative Distributions
VISUALIZING CATEGORICAL DATA THROUGH GRAPHICAL DISPLAYS
a. The Bar Chart
- visualizes a categorical variable as a series of bars.
- The length of each bar
represents either the
frequency or percentage
of values for each
category.
- Each bar is separated by
a space called a gap.
b. The Pie Chart
- is a circle broken up into slices that represent categories.
- The size of each slice of the pie varies according to the percentage in each category.
c. The Pareto Chart
- Used to portray categorical data (nominal scale)
- A vertical bar chart, where categories are shown in descending order of frequency
- A cumulative polygon is shown in the same graph
- Used to separate the “vital few” from the “trivial many”
Example: Ordered Summary Table for Causes of Incomplete ATM Transactions
d. Side By Side Chart
- Represents the data from a contingency table.
VISUALIZING NUMERICAL DATA BY USING GRAPHICAL DISPLAYS
a. Stem-and-Leaf Display
- A simple way to see how the data are distributed and where concentrations of data exist
Method:
Separate the sorted data series
- into leading digits (the stems) and
- the trailing digits (the leaves)
- A stem-and-leaf display organizes data into groups (called stems) so that the values within
each group (the leaves) branch out to the right on each row.
b. The Histogram
- A vertical bar chart of the data in a frequency distribution is called a histogram.
- In a histogram there are no gaps between adjacent bars.
- The class boundaries (or class midpoints) are shown on the horizontal axis.
- The vertical axis is either frequency, relative frequency, or percentage.
- The height of the bars represent the frequency, relative frequency, or percentage.
c. The Polygon
- A percentage polygon is formed by having the midpoint of each class represent the data in
that class and then connecting the sequence of midpoints at their respective class percentages.
- The cumulative percentage polygon, or ogive, displays the variable of interest along the X
axis, and the cumulative percentages along the Y axis.
- Useful when there are two or more groups to compare.
Frequency Polygon
Percentage Polygon
VISUALIZING TWO NUMERICAL VARIABLES BY USING GRAPHICAL DISPLAYS
a. The Scatter Plot
- are used for numerical data consisting of paired observations taken from two numerical
variables
- One variable is measured on the vertical axis and the other variable is measured on the
horizontal axis
- Scatter plots are used to examine possible relationships between two numerical variables
b. The Time Series Plot
- is used to study patterns in the values of a numeric variable over time
- Numeric variable is measured on the vertical axis and the time period is measured on the
horizontal axis
CHAPTER 2 SUMMARY
TABLES Used For Organizing CATEGORICAL Data
Summarizing Tallies the frequencies or percentages of items in a set of categories so that
Table you can see differences between categories.
Contingency Table Used to study patterns between two or more categorical variables. It
cross-tabulates or jointly tallies the responses. For two variables, one is
placed in rows and the other in columns.
TABLES Used For Organizing NUMERICAL Data
Ordered Array A sequence of data arranged in rank order from smallest to largest. It shows
the range and may help identify outliers.
Frequency A summary table that organizes data into numerically ordered classes. It
Distribution condenses raw data, provides visual interpretation, and highlights data
concentration.
Cumulative Shows cumulative frequencies or percentages up to each class, allowing you
Distribution to see how values accumulate across the range.
Visualizing CATEGORICAL Data Through Graphical Displays
The Bar Chart Visualizes a categorical variable with bars. Bar length represents frequency
or percentage, and bars are separated by gaps.
The Pie Chart A circle divided into slices, each representing a category. Slice size
corresponds to the category's percentage.
The Pareto Chart A vertical bar chart showing categories in descending order of frequency.
Includes a cumulative polygon. Separates the “vital few” from“trivial many.”
Side by Side Chart Graphically represents data from a contingency table, usually with grouped
bars for comparison across categories.
Visualizing NUMERICAL Data Through Graphical Displays
Stem-and-Leaf Displays data by splitting values into stems (leading digits) and leaves
Display (trailing digits). Shows data distribution and clustering.
The Histogram A vertical bar chart of a frequency distribution. Bars touch each other. The
horizontal axis shows class boundaries/midpoints; the vertical axis shows
frequencies or percentages.
The Polygon Uses class midpoints connected by lines to display data trends. A cumulative
polygon (ogive) uses cumulative percentages on the Y-axis.
Visualizing TWO NUMERICAL VARIABLES Through Graphical Displays
The Scatter Plot Plots paired numerical data on a graph to observe relationships between two
variables. One variable on the X-axis, the other on the Y-axis.
The Time Series Shows how a numeric variable changes over time. Time is on the X-axis, the
Plot variable’s value on the Y-axis.
CHAPTER 3: NUMERICAL DESCRIPTIVE MEASURES
Central Tendency is the extent to which all the data values group around a typical or central value.
Variation is the amount of dispersion or scattering of values
Shape is the pattern of the distribution of values from the lowest value to the highest value.
MEASURES OF CENTRAL TENDENCY
a. The Mean
- The arithmetic mean (often just called the “mean”) is the most common measure of central
tendency.
- The most common measure of central tendency
- Mean = sum of values divided by the number of values
- Affected by extreme values (outliers)
b. The Median
- In an ordered array, the median is the “middle” number (50% above, 50% below)
-
- Less sensitive than the mean to extreme values
- The location of the median when the values are in numerical order (smallest to largest):
- If the number of values is odd, the median is the middle number
- If the number of values is even, the median is the average of the two middle numbers
c. The Mode
- Value that occurs most often
- Not affected by extreme values
- Used for either numerical or categorical (nominal) data
- There may may be no mode
- There may be several modes
Example:
Which Measure to Choose?
- The mean is generally used, unless extreme value (outliers) exist.
- The median is often used, since the median is not sensitive to extreme values. For example,
median home prices may be reported for a region; it is less sensitive to outliers.
- In some situations it makes sense to report both the mean and the median.
MEASURES OF VARIATION
- Measures of variation give information on the spread or variability or dispersion of the data
values.
a. The Range
- Simplest measure of variation
- Difference between the largest and the smallest values:
Why can the Range be MISLEADING?
Does not account for how the data are distributed
Sensitive to outliers
b. The Sample Variance
- Average (approximately) of squared deviations of values from the mean
c. The Sample Standard Deviation
- Most commonly used measure of variation Formula:
- Shows variation about the mean
- Is the square root of the variance
- Has the same units as the original data
d. The Standard Deviation
Steps for Computing Standard Deviation
1. Compute the difference between
each value and the mean.
2. Square each difference.
3. Add the squared differences.
4. Divide this total by n-1 to get the
sample variance.
5. Take the square root of the sample
variance to get the sample standard
deviation.
Comparing Standard Deviations
Summary of Characteristics:
- The more the data are spread out, the greater the range, variance, and standard deviation.
- The more the data are concentrated, the smaller the range, variance, and standard deviation.
- If the values are all the same (no variation), all these measures will be zero.
- None of these measures are ever negative.
e. The Coefficient of Variation
- Measures relative variation
- Always in percentage (%)
- Shows variation relative to mean
- Can be used to compare the variability of two or more sets of data measured in different units
Comparing Coefficients of Variation
SHAPE OF A DISTRIBUTION
- Describes how data are distributed
- Two useful shape related statistics are:
Skewness
- Measures the extent to which data values are not symmetrical
Kurtosis
- Kurtosis affects the peakedness of the curve of the distribution—that is, how sharply the curve
rises approaching the center of the distribution
-
- measures how sharply the curve rises approaching the center of the distribution)
NUMERICAL DESCRIPTIVE MEASURES FOR A POPULATION (DCOVA)
- Descriptive statistics discussed previously described a sample, not the population.
- Summary measures describing a population, called parameters, are denoted with Greek letters.
- Important population parameters are the population mean, variance, and standard deviation.
a. The Mean µ
- The population mean is the sum of the values in the population divided by the population size, N
b. The Variance σ2
- Average of squared deviations of values from the mean
c. The Standard Deviation σ
- Most commonly used measure of variation
- Shows variation about the mean
- Is the square root of the population variance
- Has the same units as the original data
SAMPLE STATISTICS versus POPULATION PARAMETERS