Data Collection & Analysis
Based on Stewart Robinson (2004). Simulation: The Practice of Model
Development and Use
March 2022
Silas Maiyo (c) 2022
Introduction
• Discussed:- basic concepts behind data collection &
analysis, eg. Data definitions and data requirements for a
conceptual model; data availability; methods of
representation; statistical distributions.
• To examine:- Issues surrounding the collection and analysis
of data
Silas Maiyo (c) 2022
Data Collection & Analysis
• Data Requirements
• Data is taken to mean quantitative data, or numbers
• Numeric data are very important in simulation modelling and in
some cases large quantities of such data are required
• Data are needed on activity times, breakdown frequencies, arrival patterns and so on.
• Qualitative data – here ignore - these are non-numeric facts and beliefs about a system
that are expressed in pictures, diagrams, words or logic statements
• What is difference between DATA and INFORMATION??
• The modeller often needs to analyze data to provide information useful to the
simulation
• In any modelling exercise, data requirements can be split into THREE TYPES
(Pidd, 2009)
Silas Maiyo (c) 2022
Data Requirements - TYPES
1. The first is preliminary or contextual data.
• For the modeller and clients to develop a thorough understanding of
the problem situation some data needs to be available, for instance, a layout
diagram, basic data on process capability and beliefs about the cause of problems that
are being experienced. ( avoid large data collection exercise & detailed analysis)
• Data are very much part of the conceptual modelling process because they are
necessary for the development of the conceptual model.
2. The second data required are data for model realisation,
• Entails developing the computer model.
• In moving from the conceptual model to a computer model many data are required, for
example, detailed data on activity times and breakdowns, customer arrival patterns and
descriptions of customer types, and scheduling and processing rules.
• May require to carry out a detailed collection exercise to obtain these data - model
realisation is an output of the conceptual modelling process.
Silas Maiyo (c) 2022
Data Requirements - TYPES
3. Finally, data are required for model validation.
• Ensure that each part of the model, as well as the model as a whole, is
representing the
real world system with sufficient accuracy.
• Assuming that the real world system exists, then the obvious way to do this
is to compare the model results with data from the real system.
Silas Maiyo (c) 2022
Obtaining Data
• Having identified the data requirements the data must be obtained
• Some data are immediately available, others need to be collected.
• Three types of data encountered:
• Category A data are available either because they are known or because they have been
collected previously. E.g Bank services, Manufacturing machines
• Category B data need to be collected. Data that fall into this category often include
service times, arrival patterns, machine failure rates and repair times, and the nature of
human decision-making.
• Category C data are not available and cannot be collected. These often occur because
the real world system does not yet exist, making it impossible to observe it in operation.
E.g. Time availability is another factor, both in terms of person-time and in terms of
elapsed time available to collect meaningful data. Say, for instance, data are not
available on the repair time for a machine
Silas Maiyo (c) 2022
• Unfortunately category C data are not uncommon. Almost every
simulation study in which I have been involved has included some
category C data.
Silas Maiyo (c) 2022
Dealing with Unobtainable (Category C) Data
• There are two main ways of dealing with category C data.
• First is to estimate the data. Data may be estimated from various sources.
May be possible to obtain surrogate data from a similar system in the same,
or even another, organisation. Domain expert (staff) may provide
reasonable estimates. Uncertainty about the validity of model may reduce
the credibility.
• Second is to treat the data as an experimental factor rather than a fixed
parameter. Instead of asking what the data are, the issue is turned
around and the question is asked: what do the data need to be?
• The approach of treating category C data as an experimental factor can only
be applied when there is some control over the data in question
• Ultimately, if the data cannot be obtained and these approaches do not
suffice there are three further options.
Silas Maiyo (c) 2022
Obtaining data --- Category C
1. One is to revise the conceptual model so the need for the data is
engineered out of the model. This is not always possible.
2. The second is to change the modelling objectives so the data are no
longer needed. Of course, if the objective in question is critical, this
is not a satisfactory solution.
3. The third is to abandon the simulation study altogether, but then
the organisation is potentially left to make decisions in the absence
of any information
Silas Maiyo (c) 2022
Data Accuracy
• Albeit that data may be available (category A), it does not follow that they are
necessarily accurate.
• The source of the data should be investigated.
• Is it likely that there are errors due to the nature of the data collection exercise?
• For what purpose were the data collected?
• If it is very different from the intended use in the model, are the data still useable?
• Draw a graph of the data and look for unusual patterns or outliers.
• If the data are considered to be too inaccurate for the simulation model, then
an alternative source could be sought.
• If this is not available, then expert judgement and analysis might be used to
determine the more likely values of the data.
Silas Maiyo (c) 2022
Data Accuracy
• When collecting data (category B) it is important to ensure that the
data are as accurate as possible.
• Data collection exercises should be devised to ensure that the sample
size is adequate and as far as possible recording errors are avoided.
• What mechanisms can be put in place to monitor and avoid
inaccuracies creeping into the data collection?
• In all this, the trade-off between the cost and time of collecting data
versus the level of accuracy required for the simulation should be
borne in mind.
Silas Maiyo (c) 2022
Data Format
• Data need not just be accurate, they also need to be in the right format for the simulation.
• Time study data are aggregated to determine standard times for
activities. In a simulation the individual elements (e.g. breaks and process inefficiencies)
are modelled separately. As a result, standard time information is not always useful (in the
right format) for a simulation model.
• It is important that the modeller fully understands how the computer
model interpret the data that are input. Using the wrong interpretation for data would
yield a very different result from that intended.
• The modeller needs to know the format of the data that are being
supplied or collected and to ensure that these are appropriate for the simulation model.
• If they are not, then the data should be treated as inaccurate and actions taken to improve
the data or to find an alternative source.
• The last resort is to treat the data as category C.
Silas Maiyo (c) 2022
Representing Unpredictable Variability
• Modelling variability, especially unpredictable (or random) variability, is at the
heart of simulation modelling. Many aspects of an operations system are
subject to such variability, for instance, customer arrivals, service and
processing times, and routing decisions.---random numbers are used to
represent variability
• Modeller must determine how to represent the variability that is present in
each part of the model.
• There are basically three options available:
• Traces;
• Empirical distributions; and,
• Statistical distributions
• Also, another option for modeling unpredictable variability is bootstrapping
Silas Maiyo (c) 2022
Representing Unpredictable Variability - TRACES
• A trace is a stream of data that describes a sequence of events.
• Typically it holds data about the time at which the events occur. It may also
hold additional data about the events such as the type of part to be processed
(part arrival event) or the nature of the fault (machine breakdown event).
• The trace is read by the simulation as it runs and the events are recreated in
the model as described by the trace. The data are normally held in a data file or
a spreadsheet.
• Traces are normally obtained by collecting data from the real system.
• Automatic monitoring systems often collect detailed data about system events
and so they are a common source of trace data. Example is a call centre….
Silas Maiyo (c) 2022
Representing Unpredictable Variability – Empirical Distributions
• An empirical distribution shows the frequency with which data values, or
ranges of data values, occur and are represented by histograms, bar charts or
frequency diagrams. They are normally based on historic data.
• Indeed, empirical distributions can be formed by summarising the data in a
trace. As the simulation runs, values are sampled from empirical distributions
by using random numbers.
• If the empirical distribution represents ranges of data values, then it should be
treated as a continuous distribution so values can be sampled anywhere in the
range – See figure next slide, below…
• If the data are not in ranges, for instance, the distribution describes a fault type
on a machine, then the distribution is discrete and should be specified as such.
Most simulation software give the option to define an empirical distribution as
either continuous or discrete. Silas Maiyo (c) 2022
Silas Maiyo (c) 2022
Representing Unpredictable Variability – Statistical Distributions
• Statistical distributions are defined by some mathematical function or
probability density function (PDF). There are many standard statistical
distributions available to the simulation modeller.
• Perhaps the best known is the normal distribution that is specified by two
parameters: mean (its location) and standard deviation (its spread).
• For a given range of values of x, the area under the curve gives the probability
of obtaining that range of values. So for the normal distribution the most
likely values of x occur around the mean, with the least likely values to the far
right-hand and left-hand sides of the distribution.
Silas Maiyo (c) 2022
• The normal distribution only has limited application in simulation modelling. It
can be used, for instance, to model errors in weight
or dimension that occur in manufacturing components.
• One problem with the normal distribution is that it can easily generate negative
values, especially if the standard deviation is relatively large in comparison to
the mean
• Where data is not normally distributed, other statistical distribution
needs to be used. Are split into THREE types:
• Continuous distributions: for sampling data that can take any value across a
range
• Discrete distributions: for sampling data that can take only specific values
across a range, for instance, only integer or non-numeric values
• Approximate distributions: used in the absence of data
Silas Maiyo (c) 2022
• Continuous Distributions – this includes..
• Negative exponential - has only one parameter, the mean.
• Erlang distribution - s based on his observations of queuing in telephone systems. It
has two parameters, the mean and a positive integer k. The value of k determines the
skew of the distribution, that is, the length of the tail to the right.
• Weibull distribution - has two parameters, the shape and scale. The shape
determines the skew of the distribution. If the shape is one, then the Weibull
distribution is the same as a negative exponential distribution whose mean equals
the scale.
• log-normal distribution - Looks very similar to the Erlang and gamma distributions,
but the probability of the mode can be much higher than for those distributions.
Hence, the sample value is more likely to be around the mode. The spread defines
the skew of the distribution with higher values increasing the length of the tail to the
right
Silas Maiyo (c) 2022
Silas Maiyo (c) 2022
Silas Maiyo (c) 2022
• Discrete Distributions – this includes..
• Binomial Distribution - describes the number of successes, or failures, in a
specified number of trials. Distribution has two parameters, the number of
trials and the probability of success. E.g Rolling a DIE with 1/6 chances of
getting number six.
• Poisson Distribution – is used to represent the number of events that occur
in an interval of time, for instance, total customer arrivals in an hour. can
also be used to sample the number of items in a batch of random size, for
example, the number of boxes on a pallet. It is defined by one parameter,
the mean. The Poisson distribution is closely related to the negative
exponential distribution in that it can be used to represent arrival rates (λ),
whereas the negative exponential distribution is used to represent inter-
arrival times (1/λ).
Silas Maiyo (c) 2022
Silas Maiyo (c) 2022
• Approximate Distributions
• Are not based on theoretical underpinnings, but they provide a useful
approximation in the absence of data. As such they are useful in
providing a first pass distribution, particularly when dealing with
category C data.
• Uniform distribution, is its form, which can either be discrete or
continuous
• This distribution is useful when all that is known is the likely minimum
and maximum of a value. It might be that little is known about the size
of orders that are received by a warehouse, except for the potential
range, smallest to largest.
Silas Maiyo (c) 2022
• Traces versus Empirical Distributions versus Statistical
Distributions
• Which of the three approaches for modelling unpredictable variability
should be preferred?
• Describe Advantages and Disadvantages for each of the approaches??
• Explain Bootstrapping as a fourth approach for modeling unpredictable
variability ??
Silas Maiyo (c) 2022