0% found this document useful (0 votes)
15 views9 pages

Unit2 Study Notes

This document provides study notes for a unit on analytical techniques focusing on data sources, types, collection methods, and ethical considerations. It outlines the differences between internal and external data, primary and secondary data, and qualitative versus quantitative data, along with their respective advantages and disadvantages. Additionally, it emphasizes the importance of recognizing bias in data collection and adhering to ethical standards in research.

Uploaded by

mxolisiiy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views9 pages

Unit2 Study Notes

This document provides study notes for a unit on analytical techniques focusing on data sources, types, collection methods, and ethical considerations. It outlines the differences between internal and external data, primary and secondary data, and qualitative versus quantitative data, along with their respective advantages and disadvantages. Additionally, it emphasizes the importance of recognizing bias in data collection and adhering to ethical standards in research.

Uploaded by

mxolisiiy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MANCOSA

Bachelor of Commerce in Information and Technology Management

ANALYTICAL TECHNIQUES
UNIT 2: THE NATURE OF DATA, DATA COLLECTION AND
SOURCES
Study Notes | Compiled from Dr H. Matsongoni's Materials & Module Guide

Unit Learning Outcomes


By the end of this unit you should be able to:
• Establish a data source and determine the reliability of that source.
• Recognise the type of data being collected or presented.
• Use the appropriate technique for data collection and understand its advantages and
disadvantages.
• Advantageously combine different collection techniques.
• Identify various sources of bias in data collection and ways of preventing bias.
• Identify ethical issues in research and ensure that respondents are not harmed by
the study.

1. Data Sources (Section 2.1)


When confronted with a situation requiring data, it is often surprising how much is actually
available — in newspapers, magazines, the Internet, and internal company records. Wegner
(1999) identifies four data source categories:

Internal External Primary Secondary


Data generated inside Data from sources Data collected first- Data collected
your own organisation outside your hand by the previously for another
organisation researcher purpose

1.1 Internal Data Sources


Internal data is generated during normal business activities within the organisation. Examples
include:
• Financial data — sales vouchers, credit notes, accounts receivable.
• Production data — monthly output figures, defect rates, work-in-progress (WIP)
levels.
• Human Resource data — time sheets, staff demographics, wage and salary
schedules.
• Marketing data — monthly sales reports, advertising expenditure, customer profiles.
1.2 External Data Sources
External data comes from sources outside the organisation. Much of this is freely available
online or in business publications. The cost depends on the source.

Private Sources
• South African Chamber of Business (SACOB)
• Business Partners (previously Small Business Development Corporation)
• Industrial Development Corporation (IDC)
• Bureau of Economic Research
• Bureau of Market Research
• Bureau of Financial Analysis
• SA Labour Development Research Unit

Public Domain Sources


• Newspapers, journals, trade magazines
• Reference libraries
• Bank economic reports
• Human Sciences Research Council (HSRC)
• Council for Scientific and Industrial Research (CSIR)
• Statistics South Africa (Stats SA) — all data available at [Link]

💡 Practical Tip
New motor vehicle sales (published monthly for all NAAMSA members) are widely regarded as a
useful indicator of the state of the economy.
Stats SA is a particularly valuable free resource — virtually all its data is accessible on its website.

1.3 Primary Data Sources


Primary data is collected by the researcher directly at the point where it is generated, typically
with a specific purpose in mind.

Advantages ✓ Disadvantages ✗
Directly relevant to the problem at hand. Can be time-consuming to collect.
Greater control over data accuracy. Generally more expensive to collect.

1.4 Secondary Data Sources


Secondary data has been collected by other individuals or agencies for purposes other than
your current study. It already exists — either inside or outside the organisation.
Examples of secondary data:
• 'Aged' market research figures from a previous study.
• Previous financial statements of a company.
• An industry-wide market research report from which you extract company-specific
data.

Advantages ✓ Disadvantages ✗
Data already exists — no May not be specific or relevant to your problem.
collection effort.
Access time is relatively short. Data may be dated and hence inappropriate.
Generally less expensive to Difficult to verify accuracy or reliability.
acquire.
May not be suitable for further manipulation.

May have been 'massaged' — key omissions or extrapolations may


lead to wrong conclusions.

2. Data Types (Section 2.2)


Understanding the nature of data is necessary for two reasons (Wegner, 1999):
1. To assess data quality.
2. To select the appropriate statistical method to analyse the data.

💡 Critical Rule
An incorrect application of a statistical method to a particular data type can render findings INVALID.
Always identify your data type before choosing an analysis method.

2.1 Qualitative vs. Quantitative Data

Qualitative (Categorical) Data Quantitative (Numerical) Data


Provides labels or names for categories of items. Measures how much or how many of something.
Single observation is a word or code representing Single observation is a number representing an
a class. amount or count.
Codes are labels only — cannot be manipulated Can be meaningfully manipulated using
arithmetically. conventional arithmetic.
Examples: management level, wine preference Examples: age, distance, number of items sold,
(yes/no), gender. income.

2.2 Measurement Scales (Data Classifications)


Wegner (1999) defines four measurement scales. The first two are associated with qualitative
data; the last two with quantitative data:
Scale Data Type Key Property Example
Nominal Qualitative Categories with NO ordering or ranking. Gender, management level,
Labels only. car brand.
Ordinal Qualitative Categories WITH a meaningful T-shirt sizes (S/M/L), turnover
order/ranking, but equal spacing is not brackets (<5m / 5-10m /
guaranteed. >10m).
Interval Quantitative Ordered AND equally spaced. No true Temperature (°C), Likert scale
zero — ratios are NOT meaningful. responses (1–5).
Ratio Quantitative Ordered, equally spaced, AND has a true Age, distance, time, mass,
zero — ratios ARE meaningful. Strongest sales, income.
scale.

💡 Key Distinction: Interval vs. Ratio


Both interval and ratio scales have order and equal spacing.
ONLY ratio scale has a true zero — meaning a value of zero indicates complete absence of the
attribute.
Example: 0°C does NOT mean 'no temperature' (interval). But R0 income DOES mean 'no income'
(ratio).
Because of this, you CAN say R100,000 is twice R50,000 (ratio), but you CANNOT say 40°C is twice
20°C (interval).

2.3 Discrete vs. Continuous Data


Quantitative data is further divided into discrete and continuous:

Discrete Data Continuous Data


Can only take specific values — Can take any value within a range — no gaps between
usually whole (integer) numbers. possible values.
Usually obtained by COUNTING. Usually obtained by MEASURING.
Finite or countable number of choices. Values are real numbers — usually rounded. Boundaries
always have one more decimal place and end in 5.
Examples: number of students in a Examples: time to travel to work; tensile strength of steel;
class; cars sold in a month. speed of an aircraft.

💡 Discrete vs. Continuous — Quick Memory Aid


DISCRETE → COUNTING → Can't have 2.63 people in a room.
CONTINUOUS → MEASURING → Time, weight, temperature — these flow along a scale.
Boundary rule: e.g. a recorded value of 64 actually represents anything from 63.5 up to (but not
including) 64.5.
3. Data Collection Methods (Section 2.3)
Wegner (1999) identifies three main approaches to gathering data for statistical analyses:
3. Observation
4. Interviewing
5. Experimentation

3.1 Observation
Observation involves systematically selecting, watching, and recording the behaviour or
characteristics of living beings, objects, or phenomena. It can be used to collect both primary
and secondary data.

Types of Observation
Type Description
Participant The observer takes part in the situation being observed. Example: a doctor
hospitalised with an injury who observes hospital procedures from within.
Non- The observer watches openly or concealed but does not participate. Example:
participant 'mystery shoppers' testing whether antibiotics can be obtained without a prescription.
Open The subject knows they are being observed. Example: 'shadowing' a health worker
with their permission.
Concealed The subject is unaware of being observed, reducing bias. Example: observing brand-
choice purchase behaviour in a store.

Advantages ✓ Disadvantages ✗
Respondent is often unaware of being observed — Passive — limited opportunity to probe for
more natural behaviour, less bias. reasons or investigate further.
Produces accurate behavioural data (e.g. traffic Time-consuming — especially for small-scale
counts, quality inspections). studies of human behaviour.

3.2 Interviewing
An interview is a data-collection technique that involves oral questioning of respondents, either
individually or as a group. Data can be gathered through personal (face-to-face), telephone,
or postal/email interviews.

Degree of Flexibility
Interviews vary in how structured they are:
• High flexibility (unstructured/loosely structured): Uses a list of topics rather than fixed
questions. Useful for sensitive subjects, exploratory studies, or when the researcher
has little prior knowledge of the problem.
• Low flexibility (structured): Uses a fixed list of questions in a standard sequence with
pre-categorised answers. Appropriate when the researcher is relatively
knowledgeable, or when a large sample is involved.

Types of Interview Methods


Method Advantages ✓ Disadvantages ✗
Personal Accurate data obtained immediately; Time-consuming and expensive if trained
(Face-to- qualitative probing possible; non-verbal interviewers required.
Face) responses can be observed.
Telephone Call-backs possible; respondents feel Non-verbal responses cannot be
safer at home; cost-effective; larger observed.
sample reachable.
Postal / Mail / Best for large or geographically Very low response rates (5%–15%);
Email dispersed populations; anonymous questions must be short; no control over
responses tend to be more honest. who responds or ability to probe.

The Written Questionnaire


A written questionnaire (also called a self-administered questionnaire) is a tool where written
questions are answered by respondents in written form. It can be administered by:
• Sending questionnaires by mail with instructions and requesting mailed responses.
• Gathering respondents in one place at one time and letting them complete the form.
• Hand-delivering questionnaires and collecting them later.

💡 Critical Point on Questionnaire Design


The design of the questionnaire is critical. Poorly designed questionnaires can lead to:
• Incorrect research questions being addressed.
• Inaccurate or inappropriate data being collected.
• Response errors (people don't know, won't say, or overstate).

3.3 Experimentation
Primary data can be generated through the manipulation of variables under controlled
conditions. The researcher monitors and records the primary variable under study while
consciously controlling the effects of other influencing factors.
Examples:
• Measuring the hardness of toughened glass at various tempering furnace
temperatures.
• Measuring advertising effectiveness by manipulating the frequency and choice of
media channels.

Advantages ✓ Disadvantages ✗
Good quality data if the experiment is correctly Costly and time-consuming.
designed and executed.
Results are replicable under the same May be impossible to control all extraneous factors
conditions. — results can be distorted.

4. Bias in Data Collection and Ethical Considerations

4.1 Sources of Bias


Bias can distort results and lead to incorrect conclusions. Key sources include:
• Sampling bias — using non-random samples that don't represent the population.
• Response bias — respondents do not know, are unwilling to say, or deliberately
overstate answers.
• Observer bias — the presence of the observer affects the behaviour being studied.
• Questionnaire bias — poorly worded or leading questions that guide respondents
towards particular answers.
• Data bias — secondary data may have been 'massaged' (manipulated), with key
omissions or incorrect extrapolations.

4.2 Ethical Considerations in Research


Researchers have an ethical obligation to protect the people who participate in their studies:
• Participants must not be harmed — physically, psychologically, or socially — by the
study.
• Informed consent should be obtained before data is collected.
• Anonymity and confidentiality of respondents should be respected.
• Researchers should be transparent about the purpose of the study.
• Data should not be misrepresented or manipulated to suit a desired conclusion.

5. Activity: Data Type Classification (Worked Examples)


For each variable, identify: (i) Data type — Qualitative or Quantitative; (ii) Measurement scale
— Nominal, Ordinal, Interval, or Ratio; and (iii) if Quantitative — Discrete or Continuous.

Variable Data Type Scale D/C


Shelf life of milk Quantitative Ratio Continuous
Number of life policies issued per day Quantitative Ratio Discrete
Area of a shop floor Quantitative Ratio Continuous
Number of pages in a textbook Quantitative Ratio Discrete
Flavours in Dogmor food chunks Qualitative Nominal N/A
Wood types for making a desk Qualitative Nominal N/A
Shoe size categories Qualitative Ordinal N/A
Voltage produced by a generator Quantitative Ratio Continuous
Car types in the Mercedes range Qualitative Nominal N/A
Yes/No/Sometimes — 'Do you drink Gin?' Qualitative Nominal N/A
Number of loaves sold daily by a bakery Quantitative Ratio Discrete
Income per day of a bakery Quantitative Ratio Continuous
Monthly birth-rate at a maternity hospital Quantitative Ratio Discrete
Mass of babies at birth Quantitative Ratio Continuous
Daily distance travelled by a courier truck Quantitative Ratio Continuous
Names of teams in a cricket league Qualitative Nominal N/A

6. Unit Summary

Topic Core Idea


Data Sources Internal, External, Primary, Secondary — each has its own advantages,
disadvantages, and use cases.
Primary Data Collected first-hand; highly relevant but costly and time-consuming.
Secondary Data Pre-existing data; quick and cheap but may be outdated, biased, or irrelevant.
Qualitative Data Categorical labels or codes. Cannot be arithmetically manipulated.
Quantitative Data Numerical — measures how much or how many. Can be arithmetically
manipulated.
Nominal Scale Categories only — no order. E.g. gender, brand names.
Ordinal Scale Ordered categories — but gaps are not equal. E.g. small/medium/large.
Interval Scale Equal spacing, no true zero. Ratios meaningless. E.g. temperature, Likert scale.
Ratio Scale Equal spacing + true zero. Ratios meaningful. Strongest scale. E.g. income, age.
Discrete Data Counted — only specific values possible. E.g. number of students.
Continuous Data Measured — any value in a range. E.g. weight, time, distance.
Collection Observation, Interviewing (personal/telephone/postal), and Experimentation.
Methods
7. Self-Test Questions
Use these questions to check your understanding of Unit 2:

6. Distinguish between internal and external data sources. Provide two examples of
each.
7. What are the key differences between primary and secondary data? What are the
main risks of relying on secondary data?
8. Explain the difference between qualitative and quantitative data. Give two examples
of each.
9. A researcher conducts a customer satisfaction survey using a scale of 1 to 5. What
measurement scale is being used? Explain your answer.
10. Explain the difference between interval-scaled and ratio-scaled data. Why does this
distinction matter statistically?
11. What is the difference between discrete and continuous data? Classify the following:
(a) number of defects per batch, (b) temperature of a furnace.
12. Describe the three main data collection methods. What are the advantages and
disadvantages of each?
13. What is a population frame? Why is questionnaire design so critical in interview-
based research?
14. A company wants to study employee satisfaction. It emails a questionnaire to all
staff. Identify: (a) the data source type, (b) the collection method, and (c) one
potential source of bias.
15. Why is it important to identify the measurement scale before selecting a statistical
analysis method?

Prescribed Reading
• Anderson, D., et al., (2020), Statistics For Business And Economics, 6th Edition,
Cengage Learning, South Africa.

Recommended Reading
• Keller, G., and Gaciu, N., (2020), Statistics For Management And Economics, 2nd
Edition, Cengage Learning, South Africa.

End of Unit 2 Study Notes

You might also like