Week 1 INTRODUCTION TO FORECASTING 1 Introduction Forecasting is the
art and science of predicting future events or trends. It’s a systematic
process that analyzes historical data, identifies patterns, and uses statistical
models to project these patterns into the future. The goal is to provide a best
estimate of what will happen in each timeframe, which can be used to inform
strategic planning and operational decisions. It’s widely used in various fields
such as finance, economics, weather prediction, and supply chain
management. 1.1 Importance of Forecasting Forecasting is vital for preparing
for the future. It allows businesses to anticipate market trends, consumer
behavior, and economic conditions. In environmental science, it helps predict
weather patterns and climate change impacts. In finance, it guides
investment decisions and risk assessment. Forecasting is essential for
organizational departments to develop and implement their strategies
effectively. The finance department relies on forecasts for projecting cash
flows and determining capital needs. The human resources department
utilizes forecasts to predict staffing requirements. Similarly, the production
department depends on forecasts to organize production schedules, labor,
material needs, and inventory management. In the context of a university,
departments that would require forecasting include admissions for student
enrollment predictions, finance for budgeting and financial planning, facilities
management for maintenance and development, and aca demic
departments for course demand and staffing. Without forecasting,
organizations would struggle to adapt to change and could miss
opportunities or fail to mitigate threats. 1.2 Why Forecast? Forecasts are
inherently imperfect. Despite this, we engage in forecasting across var ious
domains weather, traffic, stock markets, and even our company’s status.
Nearly every business endeavor relies on forecasting, although not all
methods are complex. Ul timately, well-informed guesses about the future
are more valuable for planning than having no forecasts at all. The following
examples demonstrate how forecasting is a versatile tool that can be applied
across various sectors to aid in planning and decision-making. 1 Example
1.1. (a) Environmental Conservation: In the context of environmental
conservation, forecasting can be used to predict the impact of human
activities on wildlife populations. For instance, if a forecast based on habitat
loss and climate change data suggests a significant decline in the pop ulation
of a particular species, conservationists can implement measures to protect
the species’ habitat and mitigate the effects of climate change, thereby
preventing or slowing down the population decline. Forecasting in the
environmental context is essential for preparing for natural events and
mitigating their impacts on our lives and activities. (b) Weather Forecasting:
Meteorologists use various models and data to predict future weather
conditions, such as temperature, precipitation, and wind speed. For instance,
if a forecast predicts heavy rainfall in a region, this information can be used
by farmers to plan their crop irrigation schedules, by municipalities to
prepare for potential flooding, and by individuals to plan their outdoor
activities. Accurate weather forecasts help in minimizing damage and ensure
safety by allowing people and organizations to take preventive measures. (c)
Financial Forecasting: Financial analysts use historical data, current market
trends, and economic indica tors to predict future financial conditions and
stock market trends. For example, a financial forecast might predict that the
stock price of a renewable energy company will rise due to increasing
demand for clean energy solutions. Investors can use this information to
make informed decisions about buying or selling stocks. (d) Supply Chain
Forecasting: In supply chain management, forecasting is used to predict
future demand for prod ucts. This helps companies manage inventory levels,
plan production schedules, and optimize logistics. For instance, if a forecast
indicates a surge in demand for electric vehicles, a car manufacturer can
increase production in advance to meet the expected increase in sales, while
suppliers can ensure they have enough raw materials to sup port this
production. 2 Decision-making and planning In the realm of business, leaders
are constantly faced with events that can either be in fluenced by their
actions (internal) or are beyond their direct control (external). Internal events
include all decisions that a company makes within its operations, from
product development to marketing campaigns. These are areas where the
company has agency and can decide the best course of action. External
events, however, such as shifts in the global economy, changes in consumer
behavior, or regulatory changes, require a different approach. Here,
forecasting is key—it’s the tool businesses use to predict these changes and
prepare for them. By analyzing trends and data, companies can make
educated guesses about what the future holds and plan accordingly. Planning
is the overarching activity that encompasses both forecasting and decision
making. It’s the blueprint that guides a company through both the
predictable and the 2 uncertain. Effective planning means aligning the
company’s internal capabilities with the external environment to achieve
strategic goals. • Decision-Making: Involves choices within the company’s
control. • Forecasting: Predicts external market and environmental changes.
• Planning: Combines forecasting and decision-making to set a strategic
direction. • Success: Depends on managing both internal and external
factors effectively. Decision-making and planning are integral parts of the
forecasting process. They involve using forecasts to inform decisions and
create plans that guide future actions. Here’s an example that illustrates this
relationship: Example 2.1. Urban Development Planning: A city council is
considering the development of a new residential area. Before making any
decisions, they use forecasting to predict future population growth based on
current trends, birth rates, immigration, and other factors. The forecast
indicates that the city’s population will increase significantly over the next
decade. With this information, the decision-making process begins. The
council must decide how to accommodate the growing population. They
consider various options, such as expanding existing residential zones,
building new housing de velopments, or a combination of both. Planning
comes into play as they create a com prehensive urban development plan.
This plan includes timelines for construction, budget allocations,
infrastructure upgrades, and environmental impact assessments. It
integrates the forecasted population growth with the council’s decisions on
how to best meet the city’s future needs. In this example, forecasting
provided the data needed to make informed decisions, and planning ensured
that those decisions were organized into a coherent strategy for action. This
demonstrates how forecasting, decision-making, and planning work together
to ad dress future challenges and opportunities 3 Robust Areas Here are
examples for each of the robust areas in forecasting: (a) Scheduling: In
project management, a schedule forecast is used to estimate the time
required to complete a project. For instance, a construction company might
use forecasting to schedule the phases of building a new housing
development, ensuring that resources like labor and materials are available
when needed. (b) Acquiring Resources: Resource forecasting helps
businesses anticipate the need for resources and plan for their acquisition.
For example, a tech company might forecast the need for additional software
engineers and begin the recruitment process in advance to ensure they have
the necessary personnel to meet project deadlines. 3 (c) Determining
Resource Requirements: This involves forecasting to decide on the long-term
resources needed by an organization. A manufacturing firm, for ex ample,
might use forecasting to determine the amount of raw materials required to
meet production targets for the upcoming year, considering market demand
and supply chain constraints. These examples illustrate how forecasting is
applied in different areas to ensure that resources are efficiently utilized and
organizational goals are met. 3.1 Applications of robust areas in forecasting
The applications of robust areas in forecasting are diverse and can be found
across various industries and sectors. Here are some applications: (i)
Financial Sector: Banks and financial institutions use robust forecasting to
predict market trends, assess risks, and make investment decisions. For
example, a bank might use robust forecasting models to determine the
likelihood of loan defaults based on economic indicators. (ii) Supply Chain
Management: Companies apply robust forecasting to manage inventory
levels, optimize supply chain operations, and reduce costs. A retailer, for
instance, could forecast seasonal demand to ensure optimal stock levels
without overstocking or stockouts. (iii) Healthcare: Robust forecasting aids in
predicting patient admissions, managing staff schedules, and ensuring the
availability of necessary medical supplies. A hos pital may use forecasting to
anticipate the influx of patients during flu season and prepare accordingly.
(iv) Energy Sector: Energy companies forecast demand and supply to plan
production and distribution. An electric utility company might use robust
forecasting to predict peak electricity usage times and adjust generation and
grid operations. (v) Public Policy: Governments use forecasting for urban
planning, infrastructure development, and public services. For example, a
city council might forecast popu lation growth to plan for housing, schools,
and transportation needs. These applications demonstrate the importance of
robust forecasting in decision-making processes, ensuring that organizations
can respond effectively to future demands and challenges. 4 An overview of
forecasting techniques Forecasting situations vary widely in their time
horizons, factors determining actual out comes, types of data patterns, and
many other aspects. The following figures show graphs of four variables for
which forecasts might be required. 4 Based on the description of the graphs
provided, here are a few different trends that can be generated: (a)
Australian Monthly Electricity Production: This graph likely shows an up ward
trend, indicating increasing electricity production over time. This could be
due to factors such as economic growth, population increase, or the
expansion of industrial activities. (b) U.S. Treasury Bill Contracts: The
fluctuations in this graph suggest a volatile trend, with peaks and troughs
corresponding to changes in interest rates, investor sentiment, or
government borrowing. (c) Sales of Product C: The sales graph with
noticeable peaks and troughs could indicate a seasonal trend, where sales
rise during certain times of the year and fall during others, possibly due to
consumer buying habits or promotional activities. (d) Australian Clay Brick
Production: The ups and downs in this graph might reflect a cyclical trend,
with periods of high production followed by periods of low production. This
could be influenced by the construction industry’s demand cycles, housing
market trends, or economic conditions. These trends provide insights into the
variables’ behavior over time and can be used for forecasting future values
or making strategic decisions. Trend analysis is a valuable tool 5 for
identifying patterns and guiding actions based on historical data. Few more
examples of how forecasting graphs are used to visualize data and make
predictions based on historical trends and patterns. (i) Stock Market
Forecasting: Graphs in this area often show the closing price of a stock each
day. They typically feature a line graph with time on the x-axis and stock
price on the y-axis, showing trends over time. (ii) Retail Sales Forecasting:
These graphs display product sales in units sold each day for a store. They
can show seasonal trends, promotions, and other events that affect sales.
(iii) Unemployment Rate Forecasting: This involves forecasting
unemployment for a state each quarter. The graphs usually show the
unemployment rate over time and can help in policy-making and economic
planning. (iv) Gasoline Price Forecasting: Graphs here forecast the average
price of gasoline each day. They can be used by consumers and businesses
alike to make informed decisions about travel and transportation costs. 5
Categories- Forecasting Methods 5.1 Quantitative Methods Definition 5.1.
Quantitative forecasting methods rely on numerical data and mathe matical
models to make predictions about future outcomes. These methods are
based on historical data analysis, statistical techniques, and mathematical
algorithms. Characteristics: • Utilize numerical data. • Apply statistical and
mathematical models. • Focus on objective analysis. • Examples include
time series analysis, regression analysis, and machine learning algorithms.
Quantitative Method- Time Series Forecasting: Time series forecasting
involves analyzing historical data to predict future values based on patterns
observed in the data over time. This method assumes that the future
behavior of a variable can be inferred from its past behavior. Example 5.1.
Predicting the continuation of historical patterns such as the growth in sales
for a retail store. By analyzing sales data from previous years, trends and
seasonality can be identified to forecast future sales volumes accurately 6
5.2 Qualitative Methods Definition 5.2. Qualitative forecasting methods
involve subjective judgment, expert opin ions, and non-numeric information
to make predictions when quantitative data is limited, unreliable, or
insufficient. These methods rely on qualitative data such as expert opinions,
market research, and scenario analysis. Characteristics: • Depend on
subjective judgment and expert opinion. • Utilize qualitative data and
insights. • Often used when quantitative data is lacking. • Examples include
the Delphi method, expert surveys, and scenario analysis. Quantitative
Method- Explanatory Modeling: Explanatory modeling seeks to understand
how explanatory variables, such as prices and advertising expenditures, influ
ence the variable being forecasted (e.g., sales). This method typically
involves building regression models that quantify the relationship between
the explanatory variables and the target variable. Example 5.2.
Understanding how changes in advertising spending affect sales of a par
ticular product. By collecting data on advertising expenditures and
corresponding sales figures over time, a regression model can be developed
to estimate the impact of advertis ing on sales. 5.3 Unpredictable Methods
Definition 5.3. Unpredictable forecasting methods deal with situations where
the future outcomes are highly uncertain, and traditional forecasting
techniques may not apply. These methods are used when events or
phenomena are highly unpredictable, rare, or subject to significant
uncertainty. Characteristics: • Address highly uncertain or unpredictable
events. • Often involves scenario planning, wild card analysis, and sensitivity
analysis. • Focus on understanding potential outcomes and their
implications. • Examples include black swan events, extreme value theory,
and sensitivity analysis in decision-making. 5.3.1 Unpredictable Forecasting
Methods (i) Scenario Planning: When little or no information is available,
scenario planning involves constructing multiple plausible scenarios of the
future based on different assumptions and narra tives. Rather than making
specific predictions, scenario planning helps to explore a range of possible
outcomes and their implications. 7 Example 5.3. Predicting the effects of
interplanetary travel. Scenario planning might involve envisioning various
scenarios, such as successful colonization of other planets, advancements in
space tourism, or unexpected challenges in space explo ration, to
understand the potential impacts on society and technology. (ii) Wild Cards
Analysis: Wild cards analysis involves identifying low-probability, high-impact
events that could significantly alter the future landscape. While these events
are difficult to predict, analyzing their potential consequences can help
organizations prepare for unexpected developments. Example 5.4. Predicting
the discovery of a new, very cheap form of energy that produces no
pollution. While the timing and nature of such a discovery are uncertain,
conducting wild cards analysis can help anticipate its potential disruptive
effects on energy markets, geopolitics, and environmental sustainability. (iii)
Sensitivity Analysis: Sensitivity analysis involves examining how changes in
one or more variables can impact the outcome of a forecast or model. While
it doesn’t predict specific fu ture events, it helps to understand the range of
possible outcomes under different scenarios. Example 5.5. Forecasting how a
large increase in oil prices will affect the consump tion of oil products. By
conducting sensitivity analysis, one can assess how changes in oil prices
may influence consumer behavior, economic growth, and energy policy,
helping to inform decision-making and risk management strategies. These
methods provide various approaches to forecasting under different levels of
data availability and predictability, enabling organizations to make informed
decisions in un certain environments. 5.4 Other Forecasting Methods (i)
Machine Learning Forecasting: Machine learning techniques, such as neural
networks, random forests, and gradient boosting, are used to forecast future
outcomes by learning patterns from histori cal data. They are applied in
various domains, including finance, healthcare, and weather prediction.
Example 5.6. Predicting stock prices using a recurrent neural network (RNN)
trained on historical price data, trading volume, and other market indicators.
The model learns to identify patterns and make predictions based on past
market behav ior. (ii) Ensemble Forecasting: Ensemble forecasting combines
predictions from multiple models to improve accu racy and robustness. It’s
widely used in weather forecasting, risk management, and financial
modeling. 8 Example 5.7. Predicting the likelihood of a hurricane making
landfall by combining forecasts from different meteorological models, each of
which provides insights into different aspects of the storm’s behavior.
Ensemble methods reduce uncertainty and enhance the reliability of
forecasts. (iii) Bayesian Forecasting: Bayesian forecasting uses Bayesian
statistics to update beliefs about future events based on prior knowledge and
observed data. It’s applied in various fields, including marketing analytics,
healthcare, and environmental modeling. Example 5.8. Forecasting patient
readmission rates in hospitals by incorporating prior knowledge about
patient demographics, medical history, and treatment out comes into a
Bayesian model. The forecast adapts as new data becomes available,
improving accuracy over time. Each type of forecasting method has its own
strengths and weaknesses, and the choice of method depends on factors
such as data availability, the level of uncertainty, and the nature of the
forecasting problem. By selecting the appropriate method and leveraging
available data and expertise, organizations can make more informed
decisions and plan for the future effectively. An additional dimension for
classifying quantitative forecasting methods is to consider the underlying
model involved. There are two major types of forecasting models: time series
and explanatory models. 6 Explanatory Models An explanatory model is a
type of scientific model that explains the processes underlying a particular
phenomenon. It often involves identifying and understanding the cause-and
effect relationships between different variables. An explanatory model is a
type of conceptual or mathematical framework used to understand and
describe the underlying mechanisms or relationships between variables that
influence a particular phenomenon. These models aim to provide insights
into the cause-and-effect relationships governing the observed outcomes or
behaviors. Explanatory models are often employed in scientific research,
engineering, economics, social sciences, and other fields to elucidate the
mechanisms driving complex systems or phenomena. Explanatory models
assume that the variable to be forecasted exhibits an explanatory
relationship with one or more independent variables. For example, GNP =
f(monetary and fiscal policies, inflation, capital spending, imports, exports,
error). Notice that the relationship is not exact. There will always be changes
in GNP that cannot be accounted for by the variables in the model, and thus
some part of GNP changes will remain unpredictable. 9 6.1 Key features of
explanatory models (i) Identification of Variables: Explanatory models
identify relevant variables or factors that are believed to influence the
phenomenon under study. These variables may include independent
variables (causes or predictors) and dependent variables (outcomes or
responses). (ii) Establishment of Relationships: Explanatory models describe
the relationships between the identified variables, often in the form of
mathematical equations, graph ical representations, or qualitative
frameworks. These relationships can be deter ministic or probabilistic and
may involve linear or nonlinear interactions. (iii) Explanation of Mechanisms:
Explanatory models seek to explain the underly ing mechanisms or
processes that drive the observed behavior or outcomes. They provide
insights into how changes in one variable lead to changes in others and offer
explanations for observed patterns or trends. (iv) Predictive Capability: While
explanatory models primarily focus on understand ing the mechanisms
behind a phenomenon, they may also have predictive capabil ities. By
extrapolating the identified relationships, explanatory models can make
predictions about future behavior or outcomes under different conditions. 6.2
Examples of explanatory models (a) Economic Growth • An explanatory
model in economics aims to dissect the factors contributing to economic
growth within a country. It typically involves analyzing various economic
indicators such as investment, consumption, government spending, exports,
imports, inflation rates, and employment levels. • Byexaminingthecause-
and-effect relationships between these variables, economists can develop
models that explain how changes in one variable affect others and ultimately
influence the overall growth of the economy. • For example, an economist
might use regression analysis to quantify the impact of investment,
government spending, and exports on gross national product (GNP) growth
over time. The resulting explanatory model can help policy makers make
informed decisions about economic policies and interventions. (b) Weather
Forecasting • Meteorologists utilize explanatory models to predict weather
patterns by ex amining the interactions between various atmospheric
variables such as tem perature, humidity, air pressure, wind speed, and
cloud cover. • These models incorporate physical principles, mathematical
equations, and observational data to simulate the behavior of the
atmosphere and forecast future weather conditions. • For instance,
numerical weather prediction models use complex algorithms to simulate
atmospheric processes and predict weather phenomena such as rain fall,
thunderstorms, hurricanes, and temperature changes. By understanding 10
the relationships between different meteorological variables, meteorologists
can provide accurate forecasts to the public and help mitigate the impact of
severe weather events. (c) Healthcare (Epidemiology) • In epidemiology,
explanatory models are employed to investigate the trans mission dynamics
of infectious diseases and understand the factors influencing disease spread
within populations. • Epidemiologists analyze data on infection rates,
transmission routes, popu lation demographics, vaccination coverage,
environmental factors, and social behaviors to develop explanatory models
of disease transmission. • For example, epidemiological models such as
compartmental models (e.g., SIR model) and network models are used to
simulate the spread of diseases like COVID-19, influenza, and HIV/AIDS.
These models help policymakers im plement effective public health
interventions, such as vaccination campaigns, social distancing measures,
and contact tracing efforts, to control disease out breaks and minimize
morbidity and mortality. In summary, explanatory models provide valuable
insights into the underlying mechanisms driving complex phenomena,
enabling scientists, policymakers, and practitioners to make evidence-based
decisions and address societal challenges effectively. 7 Time series models
Time series models are statistical techniques used in forecasting to predict
future trends based solely on historical data of a variable or phenomenon
over time. Unlike explanatory models, which aim to understand the
underlying factors influencing the behavior of the system, time series models
treat the system as a ”black box.” This means that they do not attempt to
uncover or analyze the causal relationships between variables. Instead, they
focus on identifying patterns and trends within the historical data series and
extrapolating those patterns into the future. 7.1 Key features of time series
models 1. Historical Data Analysis: Time series models analyze past
observations of a variable collected at regular intervals over time. These
observations form a chrono logical sequence, allowing for the identification
of patterns, trends, and seasonality within the data. 2. Pattern Discovery: The
primary objective of time series modeling is to discover the underlying
patterns and structures present in the historical data series. This may involve
detecting trends (long-term movements), seasonality (repeating patterns),
cyclicality (regular fluctuations), and irregularities (random variations or
noise). 3. Forecasting: Once the patterns within the historical data have been
identified, time series models use this information to make forecasts about
future values of the variable of interest. Forecasts are typically made by
extrapolating the identified patterns into the future, assuming that historical
trends will continue. 11 4. Modeling Techniques: Various statistical
techniques are used in time series mod eling, including autoregressive
integrated moving average (ARIMA), exponential smoothing methods (e.g.,
Holt-Winters), and more advanced machine learning al gorithms such as
recurrent neural networks (RNNs) or Long Short-Term Memory (LSTM)
networks. 7.2 Examples of time series forecasting applications • Predicting
future stock prices based on historical price movements. • Forecasting
monthly sales volumes for a retail store based on past sales data. •
Estimating future demand for a product or service based on historical
consumption patterns. • Projecting quarterly GDP growth rates for an
economy using historical economic indicators. Overall, time series models
provide a powerful tool for forecasting future trends and making predictions
based solely on historical data, without the need to understand the
underlying causal mechanisms driving the system. They are widely used in
various fields, including finance, economics, marketing, and meteorology. 8
Steps involved in Forecasting There are five basic steps in any forecasting
task for which quantitative data are available. 1. Problem Definition: • Define
the forecasting problem by understanding its purpose, stakeholders, and
organizational context. • Gather information from relevant personnel and
systems. • Understand how forecasts will be used and what outcomes are
desired. 2. Gathering Information: • Collect both statistical data and expert
judgment from key personnel. • Gather historical data related to the
forecasting problem. • Data may include numerical data, reports, and
expertise from individuals in volved in the process. 3. Preliminary
(Exploratory) Analysis: • Visualize and analyze historical data to understand
patterns, trends, and anoma lies. • Conduct descriptive statistics and
exploratory data analysis to identify impor tant characteristics of the data. •
Determine the presence of trends, seasonality, business cycles, and outliers.
12 4. Choosing and Fitting Models: • Select appropriate quantitative
forecasting models based on preliminary anal ysis. • Fit chosen models to
historical data, adjusting model parameters using known data. • Consider
various forecasting techniques such as exponential smoothing, regres sion,
Box-Jenkins ARIMA models, etc. • For long-term forecasting, use less formal
approaches such as identifying mega trends, using analogies, and
constructing scenarios. 5. Using and Evaluating a Forecasting Model: • Utilize
the selected forecasting model to generate forecasts for future periods. •
Evaluate the performance of the model using future data. • Assess the
accuracy of forecasts and compare them to actual outcomes. • Consider the
broader impact of forecasts on organizational decision-making and actions. •
Incorporate new information from forecasts into management decisions to en
hance the likelihood of favorable outcomes. Step 1: Problem Definition The
bakery needs to forecast the demand for bread to manage baking schedules
and ingredient purchases efficiently. The forecasts will be used by the bakery
manager to plan daily production and by the purchasing department to order
flour and other supplies. Example 8.1. The bakery manager says, “We need
to predict how many loaves of each type of bread we’ll sell each day to
minimize waste and ensure we meet customer demand.” Step 2: Gathering
Information The bakery collects sales data from the past two years, noting
any events that may have affected sales, like holidays or promotions.
Example 8.2. The bakery reviews daily sales records, noting that sales
increase on weekends and during the local farmers’ market. Step 3:
Preliminary Analysis The bakery analyzes the sales data, looks for patterns
and trends, and calculates basic statistics to understand the variability in
daily sales. Example 8.3. The bakery observes a consistent increase in
demand for whole wheat bread during January, possibly due to New Year’s
resolutions. Step 4: Choosing and Fitting Models The retailer selects several
forecasting models suitable for retail sales, such as expo nential smoothing
to account for trends and seasonality, and regression models to include
factors like marketing spend and economic indicators. 13 Example 8.4. The
retailer examines past sales data and fits a model that accounts for seasonal
trends, promotional events, and economic conditions, helping to predict
future sales volumes. Step 5: Using and Evaluating a Forecasting Model The
retailer uses the chosen model to forecast the upcoming season’s sales and
continuously evaluates the model’s accuracy by comparing predicted sales
against actual sales figures. Example 8.5. If the model predicts a downturn in
sales, the retailer may plan pro motions or discounts to boost demand.
Conversely, if a sales increase is forecasted, they might optimize inventory
levels to meet the expected rise in customer purchases. 9 Measures of
Forecasting 9.1 Univariate Statistics Univariate statistics involve the analysis
of a single variable. The primary goal is to describe and summarize the
data’s distribution and to identify patterns within it. Here are some key
aspects of univariate analysis: • Measures of Central Tendency: These include
the mean (average), median (middle value), and mode (most frequent
value), which help indicate the central point of the data’s distribution. •
Measures of Dispersion: These include the range (difference between the
highest and lowest values), variance (average of the squared differences
from the mean), and standard deviation (square root of the variance), which
describe the spread of the data. • Frequency Distributions: This is a summary
of how often each value occurs in the dataset, often represented in a table. •
Charts and Graphs: Visual tools like histograms, boxplots, and pie charts are
used to illustrate the data’s distribution. Example 9.1. Imagine a school
wants to analyze the test scores of a class in mathemat ics. The test scores
out of 100 for 10 students are as follows: 78, 85, 92, 67, 75, 80, 84, 90, 77,
and 82. Univariate analysis of this data might include: • Mean: The average
score, which is (78 + 85 + 92 + 67 + 75 + 80 + 84 + 90 + 77 + 82)/10 =
81) • Median: The middle score when arranged in order, is 80.5 (the average
of the 5th and 6th scores, which are 78 and 83 when sorted). • Mode: The
most frequent score, which in this case, there isn’t one as all scores are
unique. 14 • Range: The difference between the highest and lowest scores,
which is ( 92- 67 = 25 ). • Standard Deviation: A calculation that shows the
average amount by which scores deviate from the mean. In this case, let’s
say it’s approximately 7.5. To find the standard deviation for the given
example, we’ll follow these steps: Step 1: Calculate the mean (average) of
the scores. Step 2: Subtract the mean from each score to find the deviation
of each score from the mean. Step 3: Square each of these deviations. Step
4: Calculate the mean of these squared deviations. Step 5: Take the square
root of this mean to get the standard deviation. This analysis gives the
school a clear picture of the overall performance of the class in the
mathematics test. It shows that the average score is 81, with a moderate
spread of scores around this average. 9.2 Bivariate statistics Bivariate
statistics involve the analysis of two variables to determine the relationship
be tween them. It’s used to find out if there is a connection between these
variables, which are often denoted as X and Y. For example, a researcher
might study the relationship between students’ heights (X) and their grades
(Y). If the data shows that taller students tend to have higher grades, there
might be a positive correlation between height and grades. Example 9.2.
Bivariate data Student Height (X) Grade (Y) 1 160 cm 90 2 155 cm 85 3 170
cm 95 In this table, each student’s height and grade are paired, and the
analysis might reveal patterns or relationships between the two variables.
Bivariate analysis can be visualized using scatter plots, where one variable is
plotted on the x-axis and the other on the y-axis. This helps in identifying the
type of relationship, whether it’s positive, negative, or non-existent. For
instance, if we plot the above data, we might see a trend line that shows an
upward trajectory, suggesting a positive relationship between height and
grades. Bivariate analysis is not limited to numerical data; it can also involve
categorical data. For example, analyzing the relationship between gender (X)
and choice of major (Y) in college students is also a form of bivariate
analysis. 15 In summary, bivariate statistics is a powerful tool for
understanding the connections between two different variables and can be
applied in various fields such as medicine, economics, and education. 9.3
Autocorrelation: Autocorrelation, also known as serial correlation or lagged
correlation, is a statistical measure that assesses the degree of similarity
between a given time series and a lagged version of itself over successive
time intervals. It’s used to identify patterns in data that occur over time. For
example, consider a time series of daily temperatures. If the temperature is
high today, autocorrelation would measure how likely it is that the
temperature will also be high tomorrow, or the next day, and so on. This is
done by comparing the current value of the time series with its past values.
Example 9.3. Autocorrelation in a time series: Day Temperature Monday 30°C
Tuesday 31°C Wednesday 29°C Thursday 28°C Friday 27°C If we find that
higher temperatures are followed by higher temperatures and lower by
lower, we have positive autocorrelation. If the opposite is true, we have
negative autocor relation. In financial markets, autocorrelation can be
observed in stock prices. If a stock’s price increases over several days, a
positive autocorrelation might suggest that the price is likely to continue
increasing in the short term. Mathematically, Let: • Xt, represents the value
of the time series X at time t. • ¯Xrresent the mean of the time series X. • n
represents the total number of observations in the time series. The formula
for the autocorrelation function, ρ(k) is given by: n ρ(k) = t=k+1 (Xt − ¯X)
(Xt−k − ¯ X) n t=1 (Xt − ¯X) 2 where: • The numerator is the sum of the
product of the deviations of X from its mean at time t and time t+k, where k
represents the lag. 16 • The denominator is the sum of the squared
deviations of X from its mean over all observations. • This formula measures
the linear dependency between observations at a lag of k time units.
Autocorrelation is a key concept in time series analysis and is particularly
important in the identification of trends, seasonality, and forecasting. The
below explains the uni variate analysis, correlation and autocorrelation.
Correlation range: When (i) r=1. we will proceed to the forecasting model. (ii)
r=0, no need to go to the forecasting model 9.4 Measuring Forecast Accuracy
Measuring forecast accuracy in many instances, the word “accuracy” refers
to “goodness of fit,” which in turn refers to how well the forecasting model
can reproduce the data that are already known. To the consumer of
forecasts, it is the accuracy of future forecasts that is most important. This
text explains that forecast accuracy is about how well a forecasting model
can match known data, which is crucial for trusting the model’s future
predictions. Observation Yt: In the context of time series analysis, “Yt” refers
to the actual observed value at a specific time period (t). It represents the
real, measured data point at that time. For instance, if you’re looking at daily
stock prices, Yt could be the actual closing price of a stock on the day (t).
Forecast Ft: On the other hand, “Forecast Ft” is the predicted value at the
same time period (t), made by a forecasting model. This prediction is based
on previous data points and possibly other factors, depending on the
complexity of the model. Continuing with the stock ex ample, Forecast Ft
would be the price that a model predicts for the stock on the day (t), given
the stock’s historical prices. The accuracy of the forecast (Ft) is often
evaluated by comparing it to the actual ob servation (Yt). The closer the
forecast is to the actual observation, the better the model is considered to be
at predicting future values. The table provides a period-wise comparison of
actual observations (Yt )and forecasts (Ft), which can be used to assess the
model’s performance. Period Observation (Yt) Forecast (Ft) 1 138 150.25 2
137 139.50 3 152 157.25 4 136 143.50 17 For instance, in period 1, the
actual observation was 138, but the forecast was 150.25. The forecast was
off by 12.25 units, which indicates that for this period, the forecast was not
very accurate. Similarly, for other periods, we can calculate the difference
between the forecast and the observation to assess the accuracy. The goal of
any forecasting model is to minimize these differences, thereby improving
the accuracy of future forecasts. This is crucial because accurate forecasts
are essential for planning and decision-making in various fields, such as
finance, weather forecasting, inventory management, and more. 9.5 Method
of Least Squares The method of least squares is a standard approach for
finding the best-fitting curve to a given set of points by minimizing the sum
of the squares of the offsets (the residuals) of the points from the curve. In
the context of linear regression, it aims to find the line (or hyperplane in
higher dimensions) that best fits the data points. 9.6 Gauss-Markov Theorem
The Gauss-Markov theorem states that, under certain conditions (the
“classical linear regression model” assumptions), the ordinary least squares
(OLS) estimator is the Best Linear Unbiased Estimator (BLUE). This means
that among all the unbiased linear es timators, OLS has the smallest
variance, making it the most reliable for estimating the coefficients in a
linear regression model. 9.7 Classical Assumptions For the Gauss-Markov
theorem to hold, the following assumptions must be met: 1. Linearity: The
relationship between the independent variables and the dependent variable
is linear. 2. Independence: The residuals (errors) are independent of each
other. 3. Homoscedasticity: The residuals have constant variance at every
level of the independent variables. 4. No perfect multicollinearity: The
independent variables are not perfectly cor related with each other. 5. Zero
mean of residuals: The average of the residuals is zero. Example 9.4. Let’s
consider a simple linear regression model where we want to predict a
dependent variable (Y) based on one independent variable ( X ). The model
is represented as: Y =β0+β1X +ϵ where: • β0 is the intercept term (the
value of Y when X is 0). • β1 is the coefficient of the independent variable X
(the slope of the line). 18 • ϵ represents the error term, which captures the
difference between the observed and predicted values of Y. Using the method
of least squares, we calculate the estimates β0 andβ1 that minimizes the
sum of squared residuals: n min i=1 Yi −(ˆβ0 + ˆβ1Xi) 2 (1) If our model
satisfies the Gauss-Markov assumptions, then these OLS estimates β0 andβ1
are BLUE, meaning they are the most precise (have the smallest variance)
among all un biased linear estimators. In practice, this means that if we use
OLS to estimate the coefficients of our linear regression model, and if the
assumptions hold true, we can be confident that our estimates are as close
to the true population parameters as possible with the least amount of error.
9.8 Na¨ıve forecast and na¨ıve forecasting methods Naive forecasting is a
method used in time series analysis where the forecast for a future period is
assumed to be the same as the observed value in the previous period. This
approach is called “naive” because it doesn’t take into account any other
information such as trends, seasonal effects, or other factors that could
influence the forecast. There are different methods of naive forecasting,
including: • Simple Naive Forecast: Uses the last observed value as the
forecast for the next period1. • Seasonal Naive Forecast: Uses the last
observed value from the same season of the previous year as the forecast
for the current season2. • Moving Average Naive Forecast: Uses the average
of a fixed number of the most recent observations as the forecast3. Despite
its simplicity, naive forecasting can be surprisingly effective for certain
datasets, especially when the data doesn’t show strong trends or seasonal
patterns. However, it’s important to measure the accuracy of naive forecasts
using metrics like Mean Absolute Percentage Error (MAPE) or Mean Absolute
Deviation (MAD) to understand how well the forecast might perform.
Example 9.5. Simple examples for each of the naive forecast methods Simple
Naive Forecast: Suppose a shop sold 120 umbrellas in the last month (Febru
ary). Using the simple naive forecast, we would predict that the shop will also
sell 120 umbrellas in the next month (March). Seasonal Naive Forecast:
Imagine a resort that has a pattern of receiving 200 guests every December
due to the holiday season. Using the seasonal naive forecast, we would
predict that the resort will receive 200 guests next December as well. Moving
Average Naive Forecast: Let’s say a website had the following number of
visitors over the last four months: 1000 in October 1200 in November, 1500
in December, 19
and1300inJanuary.Usingamovingaveragenaiveforecastwithawindowof3month
s, wewouldcalculatetheforecast forFebruaryastheaverageof thelast
threeobservations: 1200+1500+1300 3 =1333.33
So,theforecastforthenumberofvisitorsinFebruarywouldbeapproximately1333.T
hese
examplesillustratehoweachmethodusesdifferentaspectsofpastdatatomakeafo
recast [Link], theeffectivenessof
thesemethodsdependsonthenatureof the dataandthecontextof theforecast.
10 ForecastingAccuracyMeasures •Howtomeasurethesuitabilityofaparticular
forecastingmethodforagivendata set. • Inmost forecastingsituations,
accuracy is treatedas theoverridingcriterion for selectingaforecastingmethod.
• Inmanyinstances, theword“accuracy”refersto“goodnessoffit,”whichinturn
refers tohowwell the forecastingmodel canreproducethedatathatarealready
known. •To the consumer of forecasts, it is theaccuracyof future forecasts
that ismost important. Explanation:
Forecastingaccuracymeasuresareusedtoevaluatehowwell
aforecastingmethodcan predict the futurevaluesofavariablebasedonpastdata.
Therearedifferentways to
measureforecastingaccuracy,dependingonthepurposeandthetypeofdata. IfYt
isthe actualobservationfortperiodtandFt
istheforecastforthesameperiod,thentheerror isdefinedas, et=Yt−Ft
•MeanError(ME): It istheaverageof thedifferencesbetweentheactualvalues
andthepredictedvalues. It indicateswhether the forecastingmethodtends to
[Link]
underestimatestheactualvalues,whileanegativeMEmeansthemethodoveresti
[Link]. MeanError(ME)=
1 n n i=1 et •MeanAbsoluteError(MAE)orMeanAbsoluteDeviation(MAD):This is
theaverageof theabsolutevaluesof thedifferencesbetweentheactual values
andtheforecastedvalues. It indicateshowlargetheerrorsare, regardlessof their
[Link].
MeanAbsoluteError(MAE)= 1 n n i=1 |et| 20 • MeanSquared Error (MSE): It is
the average of the squared differences between the actual values and the
predicted values. It also indicates how large the errors are, but it gives more
weight to larger errors than smaller errors. A smaller MSE means the method
is more accurate, especially for data with outliers or high variability. Mean
Squared Error (MSE) = 1 n n i=1 et2 • Percentage Error (PE): It is the ratio of
the difference between the actual value and the predicted value to the actual
value, expressed as a percentage. It indi cates how large the error is relative
to the actual value. A positive PE means the method underestimates the
actual value, while a negative PE means the method overestimates the
actual value. Percent Error (PEt) = et Yt × 100 or Yt−Ft Yt ×100 • Mean
Percentage Error (MPE): It is the average of the percentage errors. It
indicates whether the forecasting method tends to overestimate or
underestimate the actual values on average. A positive MPE means the
method underestimates the actual values on average, while a negative MPE
means the method overestimates the actual values on average. A zero MPE
means the method is unbiased on average. Mean Percentage Error (MPE) = 1
n n i=1 PEt • MeanAbsolute Percentage Error (MAPE): This is the average of
the absolute values of the percentage differences between the actual values
and the forecasted values. It indicates how large the errors are relative to the
actual values. A smaller MAPE means the method is more accurate,
especially for data with different scales or units. Mean Absolute Percentage
Error (MAPE) = 1 n n i=1 |PEt| 10.1 Accuracy and Goodness of Fit: Why is
important in fore casting? • Accuracy is important because it measures how
close the predictions or estimates of a model are to the actual values. A
model with high accuracy can provide reliable and useful information for
decision-making, while a model with low accuracy can lead to errors and
misleading conclusions. • Goodness of fit is related to accuracy because it
evaluates how well a model fits the data that are already known. A model
with high goodness of fit can reproduce the observed data with small and
unbiased differences. A model with low goodness of fit can have large and
biased differences between the observed and expected values. • Goodness
of fit can also help to assess the accuracy of future predictions or esti mates,
by testing whether the model is appropriate for the data and whether the 21
assumptions of the model are valid. A model that passes the goodness of fit
test can be more confident in its accuracy, while a model that fails the
goodness of fit test can be less confident or need to be revised. Examples •
Amazon: This uses AI-driven predictive forecasting to respond quickly to
unfore seen demand signals and increase adaptability to market fluctuations.
For instance, when toilet paper sales surged by 213% at the height of the
COVID-19 pandemic, Amazon’s models reacted quickly to the new demand
trend and adjusted their in ventory and supply chain accordingly. • Business
forecasting: This involves forecasting tools and techniques to help busi
nesses predict certain developments, such as revenue, sales, and growth.
For ex ample, a new company that started the year with few sales used a
qualitative fore casting technique to gauge their Q4 sales based on their Q3
performance and a new marketing technique. • Measuring forecast accuracy:
This is a critical metric used by businesses and organizations to evaluate the
effectiveness of their forecasting models. It helps them assess the reliability
and validity of their predictions, enabling them to make in formed decisions
based on the forecasted outcomes. For example, a retail store used
statistical techniques such as MAPE, RMSE, and MAD to measure the accu
racy of their forecasts for different products, such as slow movers and fast
movers. By comparing the values of these measures for different methods,
they were able to choose the most suitable method for each product
category and improve their replenishment process. • Error term: The
difference between the observed value and the predicted value of a variable
in a statistical model. For example, if a model predicts that the price of a
product is $50, but the actual price is $52, then the error term is $52- $50 =
$2. The error term reflects the uncertainty or randomness in the model, and
it is usually denoted by, e, ϵ, or u. • Mean error is the average of all the error
terms in a model. It is calculated by summing up all the error terms and
dividing by the number of observations. For example, if there are 10
observations with error terms of 2,-1, 3,-2, 1,-3, 4,-4, 5, and-5, then the mean
error is (2- 1 + 3- 2 + 1- 3 + 4- 4 + 5- 5) / 10 = 0. The mean error measures
the bias or accuracy of the model, and it is ideally zero or close to zero.
Example 10.1. You are a stock market analyst who wants to evaluate the
accuracy of your prediction model that estimates the closing price of a stock
based on the opening price, the volume, and the market sentiment. You have
the actual and predicted closing price values for the past 5 days in the table
below. Calculate the ME, MAE, PE, MPE, MSE, and MAPE for your model and
interpret the results. 22 Day Actual Closing Price ($) Predicted Closing Price
($) 1 50 2 49 52 3 51 54 4 53 56 5 55 58 57 ME: Mean Error (ME) = 1 n n i=1
et Step 1: Find the General Error et = Yt −Ft Yt = 50,52,54,56,58;Ft =
49,51,53,55,57;et = 1,1,1,1,1 Step 2: The sum of the observation is divided
by the number of observations = 5/5 = 1 MAE: Mean Absolute Error (MAE) =
1 n n i=1 |et| Step 1: Find the General Error et = Yt −Ft Yt =
50,52,54,56,58;Ft = 49,51,53,55,57;et = 1,1,1,1,1 Step 2: The sum of the
modulus of observations is divided by the number of observations = 5/5 = 1
MSE: Mean Squared Error (MSE) = 1 n n i=1 et2 Step 1: Find the General
Error et = Yt −Ft Yt = 50,52,54,56,58;Ft = 49,51,53,55,57;et = 1,1,1,1,1
Step 2: The sum of the squared general error is divided by the number of
observations = 5/5 = 1 PE: Mean Percentage Error (MPE) = 1 n n i=1 PEt Find
the General Error et = Yt −Ft Yt = 50,52,54,56,58;Ft = 49,51,53,55,57;et =
1,1,1,1,1 23 PE1 = 1 50× 100% = 2% PE2 = 1 52× 100% = 1.92% PE3 = 1
54× 100% = 1.85% PE4 = 1 56× 100% = 1.79% PE5 = 1 58× 100% =
1.72% MPE: Mean Percentage Error (MPE) = 1 n n i=1 PEt The sum of
Percentage error is divided by the number of observations. MPE=1 5
(2+1.92+1.85+1.79+1.72) =9.28/5 = 1.856% MAPE: Mean Absolute
Percentage Error (MAPE) = 1 n n i=1 |PEt| The sum of the modulus of PE is
divided by the number of observations. |PEt|
=2%,1.92%,1.85%,1.79%,1.72% MPAE =1 5 (2+1.92+1.85+1.79+1.72)
=9.28/5 = 1.856% 11 Theil’s U-statistic Theil’s U-statistic is a powerful tool
for assessing the accuracy of forecasting methods. Unlike the Mean Squared
Error (MSE), which treats all errors equally, Theil’s U-statistic takes into
account both the disproportionate cost of large errors and provides a relative
basis for comparison with naive forecasting approaches. Here’s a breakdown
of its key characteristics: (a) Relative Comparison: • Theil’s U-statistic allows
us to compare formal forecasting methods with naive (simple) approaches. •
It provides insight into how well a forecasting method performs relative to a
basic, straightforward prediction. (b) Weighting of Errors: • Large errors are
emphasized in Theil’s U-statistic. • By squaring the errors, it gives more
weight to significant deviations, making them stand out. 24 (c) Intuitive
Interpretation Sacrificed: • While Theil’s U-statistic is robust and informative,
it lacks the intuitive inter pretation found in simpler metrics. • Understanding
its computation and application can be challenging due to its mathematical
complexity. Now, let’s express Theil’s U-statistic mathematically. It is defined
as: n−1 U = where t=1 (FPEt+1 −APEt+1)2 n−1 t (APEt+1)2 FPEt+1 =Ft+1
−Yt Yt APEt+1 =Yt+1 −Yt Yt where: (forecast relative change) (actual
relative change) • n represents the number of observations. • Yt is the actual
value at time t. • Ft is the forecasted value at time t. In summary, Theil’s U-
statistic strikes a balance between accuracy assessment and the impact of
large errors, making it a valuable tool for evaluating forecasting models.
Keep in mind that while it may lack intuitive simplicity, its insights are
invaluable for informed decision-making in forecasting endeavors. Suitable
for best model, moderate, not good (a) Similarity to MAPE: • The
denominator of Theil’s U-statistic is analogous to the numerator, with one
key difference: Ft+1 (forecasted value) is replaced by Yt (actual value). •
This similarity to MAPE highlights that Theil’s U-statistic incorporates both
concepts of forecasting accuracy. (b) Theil’s U-Statistic: • Theil’s U-statistic
emphasizes large errors by squaring the differences between forecasted and
actual values. • It provides a relative comparison between forecasting
methods and considers the disproportionate cost of large errors. • While it
sacrifices intuitive interpretation, it helps us make informed decisions about
forecasting techniques. The U-statistic ranges can be summarized as follows:
25 (a) U =1 • Whenthe U-statistic equals 1, it implies that the naive method
(such as a sim ple average or no forecasting) performs equally well as the
evaluated forecasting technique. • In other words, there is no advantage in
using the formal forecasting method over the naive approach. (b) U 1 •
Conversely, when the U-statistic exceeds 1, there is no point in using a
formal forecasting method. • In such cases, sticking with the naive method
would yield better results. In summary, Theil’s U-statistic provides a relative
comparison between forecasting methods and emphasizes the impact of
large errors, helping us make informed decisions about which approach to
use. Example 11.1. Refer Video content 138 Period t Observation Yt Forecast
Ft Error Absolute error Squared error 1 150.25-12.25 12.25 150.06 2 136
139.50-3.50 3.50 12.25 3 152 157.25-5.25 5.25 27.56 4 127 143.50-16.50
16.50 272.25 5 151 138.00 13.00 13.00 169.00 6 140 127.50 2.50 2.50 6.25
7 142 138.25-9.75 9.75 95.06 8 153 141.50 3 3 9 Total-28.75 Solutions:
Mean Error (ME): Write your solution here... 55.75 ME=
(−12.25−3.50−5.25−16.50+13.00+2.50−9.75+3.00) 8 Mean Absolute Error
(MAE): MAE = 12.25+3.50+5.25+16.50+13.00+2.50+9.75+3.00 8 741.43
=−3.71875 =8.09375 26 MeanSquaredError(MSE):
MSE=150.06+12.25+27.56+272.25+169.00+6.25+95.06+9.00 8 =92.5425
= √ 92.5425=9.6204units. MeanAbsolutePercentageError(MAPE):
MAPE=12.25 138 +3.50 136 +5.25 152 +16.50 127 +13.00 151 +2.50 140
+9.75 142 +3.00 153 8 ×100% =7.841 Theil’sU-statistic: Step1:
findtheforecastrelativechangeandactualrelativechange. findthesquarevalue
ofthat. Ft+1−Yt Yt 2 =0.006,0.0015,0.0118,0.0105, 0.0003, 0.0219, 0.0093
Yt+1−Yt Yt 2 =0.0002,0.0138,0.0271,0.0357,0.0193, 0.0072, 0816
Step2:Findthesumvaluesof Ft+1−Yt Yt 2 and Yt+1−Yt Yt 2 Ft+1−Yt Yt 2
=0.0560 Yt+1−Yt Yt 2 =0.1849 Step3:NowapplytheTheil’sUformula
Theil’sU= 0.0560 0.1849=0.550 12 PredictionIntervals • It
isusuallydesirabletoprovidenotonlyforecastvaluesbutaccompanyinguncer
taintystatements,usuallyintheformofpredictionintervals.
•PredictionintervalsareusuallybasedontheMSEbecauseitprovidesanestimate
ofthevarianceoftheone-stepforecasterror.
•ThesquarerootoftheMSEisanestimateofthestandarddeviationoftheforecast
error. •Underthisassumption,anapproximatepredictioninterval
forthenextobservation isFn+1±Zα √MSE
•Thevalueofzdeterminesthewidthandprobabilityofthepredictioninterval.
•Forexample, z=1.96givesa95%predictioninterval. That is, the intervalhasa
probabilityof95%containingthetruevalue,asyetunknown.
Here,zisthevaluetakenfromthenormaldistributiontable. 27 12.1 Relative
frequencies in statistics A relative frequency describes how often a specific
value for a variable (data item) occurs with the total number of values for
that variable. Instead of using raw counts, relative frequencies express the
count for a particular type of event as a percentage, proportion, or fraction
relative to the total number of events. For example: • If 25% of the books Jim
read were about statistics, that’s a relative frequency. • If the football team
won 85% of its games, that’s another relative frequency. Relative frequencies
help us place specific events into a larger context and allow for com parisons
between different studies or scenarios. To calculate relative frequencies,
divide the count of a specific type of event by the total number of
observations. The formula for relative frequency is Relative Frequency =
Count of Specific Event Total Number of Events These relative frequencies
also serve as empirical probabilities, helping us understand the likelihood of
events occurring. 13 Confidence Intervals In statistical hypothesis testing,
the null hypothesis (denoted as H0) and the alterna tive hypothesis (denoted
as H1orHa) are two opposing statements about a population parameter.
These hypotheses are used to assess the validity of a claim based on sample
data. (a) Null Hypothesis (H0): The null hypothesis is a statement that
suggests there is no effect or no difference between groups or conditions. It
typically represents the status quo or a belief that there is no relationship
between variables. In other words, it states that any observed difference is
due to random variation or chance. (b) Alternative Hypothesis (H1 or Ha):
The alternative hypothesis is a statement that contradicts the null
hypothesis. It suggests that there is a significant effect, relationship, or
difference between groups or conditions in the population. Example 13.1.
Let’s consider an example in the context of forecasting: Null Hypothesis (H0):
There is no significant difference in the sales performance between the two
advertising strategies. Alternative Hypothesis (H1): There is a significant
difference in the sales perfor mance between the two advertising strategies.
28 In this example, suppose a company is testing two different advertising
strategies (Strategy A and Strategy B) to determine which one yields better
sales performance. The null hypothesis (H0) suggests that there is no
significant difference in sales perfor mance between the two strategies,
implying that any observed difference in sales could be due to random
chance. The alternative hypothesis (H1) contradicts this, suggesting that
there is indeed a significant difference in sales performance between the two
strategies. To map this into forecasting, one might collect sales data over a
specified period while employing both advertising strategies and then
conduct statistical tests to determine if there is a significant difference in the
sales generated by each strategy. If the null hypoth esis is rejected based on
the data, it implies that one advertising strategy is more effective for sales
forecasting than the other. If the null hypothesis is not rejected, it suggests
that there is insufficient evidence to conclude that one strategy is better
than the other in terms of sales forecasting. Hypothesis testing is a statistical
method used to evaluate and make decisions about a population based on
sample data. Here’s a step-by-step guide to the process: 1. State Your
Hypotheses: Formulate the null hypothesis (H0) which posits no effect or
difference, and the alternative hypothesis ((Ha)or(H1)) which suggests there
is an effect or difference. 2. Collect Data: Gather data in a manner that is
unbiased and representative of the population. 3. Perform a Statistical Test:
Choose and execute an appropriate statistical test that compares the
observed data with what would be expected under the null hy pothesis. 4.
Make a Decision: Based on the test’s p-value and the level of significance (α),
decide whether to reject or fail to reject the null hypothesis. 5. Present
Findings: Report the results of your hypothesis test, including the sta tistical
significance and what it implies about the population. Basic Terminologies •
Null Hypothesis (H0): This is the hypothesis that there is no significant
difference or effect. It’s the default position that assumes no association
between two measured phenomena. In mathematical terms, it could be
expressed as (H0 : µA = µB), where (µA) is the population mean and (µB) is a
specific value. • Alternative Hypothesis (H0) or (H1): This hypothesis is
contrary to the null hypothesis and suggests that there is a significant effect
or difference. It’s what researchers aim to support. Mathematically, it could
be expressed as (H1 : µA ̸ = µB)for a two-tailed test, or (H1 : µA > µB) or
(H1 : µA < µB) for one-tailed tests. • Level of Significance Hα: This is the
probability of rejecting the null hypothesis when it is true, also known as the
Type I error rate. Common significance levels are 0.05, 0.01, and 0.10. 29 •
Testing Measure/Test statistic: This refers to the statistical test used to evalu
ate the evidence against the null hypothesis. It involves calculating a test
statistic, such as a t-score or z-score, which is then compared to a critical
value to determine whether to reject (H0). • Inference/Conclusions: After
comparing the test statistic to the critical value, a conclusion is drawn. If the
test statistic falls into the critical region, (H0) is rejected in favor of (H1). If
not, there is not enough evidence to reject (H0), and it fails to be rejected.
These terms are fundamental to understanding and performing hypothesis
testing, which is a key method in statistical inference. Hypothesis testing
allows researchers to make decisions about the validity of claims based on
sample data. Types of Error Key concepts in hypothesis testing, particularly
focusing on the risks associated with making decisions based on statistical
tests. Here’s an explanation of the terms and con cepts: • Producer’s Risk
(Type I Error, Hα): This is the risk of rejecting a true null hypoth esis. It’s also
known as a false positive. The level of significance (α), often set at 0.05, is
the probability of committing a Type I error. • Consumer’s Risk (Type II Error,
Hβ): This is the risk of failing to reject a false null hypothesis. It’s also known
as a false negative. The power of the test (1 −β) is the probability of
correctly rejecting a false null hypothesis. • α: α is the threshold value that
determines the cut-off point for rejecting the null hypothesis. If the p-value of
the test is less than α, we reject the null hypothesis. • β: β is related to the
power of the test, which is the ability to detect an effect if there is one. A
high power means a low risk of committing a Type II error. • Normal Curve: In
the context of a normal distribution curve, α is represented by the area in the
tail(s) beyond the critical value(s) where we would reject (H0). The area
where we fail to reject (H0) is 1−α. The β region is the area where we fail to
reject H0 even though H1 is true, typically found between the critical value
and the true population parameter under (H1). To decide between the risks,
one must consider the consequences of each error. If the cost of a false
positive is high, a lower α is chosen. If a false negative carries more risk,
efforts are made to reduce β, often by increasing the sample size or effect
size. For a practical example, consider a pharmaceutical company testing a
new drug. The producer’s risk α would be the chance of incorrectly
concluding that the drug is effective when it’s not, potentially leading to
unnecessary costs and health risks. The consumer’s risk β would be the
chance of failing to recognize the drug’s effectiveness, possibly depriv ing
patients of a beneficial treatment. 30 In hypothesis testing, balancing these
risks involves choosing an appropriate level of significance and ensuring
sufficient sample size to achieve the desired power of the test. Hypothesis
testing is a fundamental aspect of statistical analysis and decision-making in
various fields, including medicine, manufacturing, and social sciences. 31