What is secondary data analysis?
Secondary data analysis refers to the process where a researcher analyses existing quantitative or
qualitative data that has already been collected and published by others. Instead of collecting
new data, the researcher uses available information from sources such as government reports,
academic studies, or online databases.
This method is useful because it saves both time and money, and it helps to avoid unnecessary
duplication of research efforts. It also gives researchers access to large-scale and high-quality
datasets that may be difficult to collect on their own. In some cases, it allows for new insights or
reinterpretation of existing data.
However, one limitation is that the researcher has no control over how the data was originally
collected, which may affect how suitable it is for the new study.
Secondary data analysis is usually contrasted with primary data analysis, where the researcher
collects and analyses their own data.
Popularly used quantitative and qualitative sources for secondary research
Data on the Internet
o Includes online databases, reports, and published research
National and provincial government and nongovernment agencies
o Provide large amounts of reliable and official data
Public and legal deposit libraries
o Store a wide range of documents, records, and publications
Higher educational institutions
o Universities produce research studies, theses, and academic publications
Commercial information sources
o Data collected by private organizations, often for business or market research
purposes
Advantages of secondary data analysis
Cost and time efficient
o Saves money and reduces the time needed to collect data
High-quality data
o Data is often collected through well-designed and reliable studies
Rigorous sampling procedures
o Strong methods are used to select participants, improving accuracy
Follow-up on non-respondents
o Procedures are often in place to improve response rates and reduce bias
Large-scale samples
o Many datasets include large numbers of participants, increasing reliability
Cross-national comparison
o Some datasets allow comparisons between different countries
Availability of many datasets
o A wide range of existing data is available for use
Opportunity for longitudinal analysis
o Data collected over time allows researchers to study trends and changes
Sub-group or subset analysis
o Researchers can focus on specific groups within the data
More time for data analysis
o Since data collection is not required, more time can be spent analysing data
Re-analysis for new interpretations
o Existing data can be re-examined to generate new insights or perspectives
Limitations of secondary data analysis
Lack of familiarity with the data
o Researchers may need time to understand the dataset and its variables
Complexity of the data
o Large datasets with many respondents and variables can be difficult to manage
and analyse
No control over data quality
o Researchers cannot influence how the data was originally collected
Absence of key variables
o Important variables may be missing, especially if they are central to the research
theory
Official statistics (Advantages)
Compared to other forms of quantitative data, such as survey data, official statistics offer several
advantages:
Already collected data
o Saves time and reduces the cost of data collection
Cross-sectional and longitudinal analysis
o Data can be analysed at a single point in time or over a period to identify trends
Cross-cultural analysis
o Allows comparisons between different countries or populations
📄 DOCUMENTS AS SOURCES OF DATA
Documents are pre-existing data sources that were not created for the current research.
They can be read or viewed, such as reports, media content, photographs, and videos.
They are stored and preserved, allowing analysis of past events and trends over time.
They are unobtrusive, meaning no direct interaction with participants is needed.
However, they must be critically evaluated for bias, accuracy, and purpose.
📄 CRITERIA TO ASSESS THE QUALITY OF
DOCUMENTS
Authenticity – Is the document genuine and from a trustworthy source? (Check the
author and origin.)
Credibility – Is the information accurate and free from bias, errors, or distortion?
Representativeness – Is the document typical of others like it, or is it unusual or biased
in some way?
Meaning – Is the content clear, understandable, and interpretable by the researcher?
📄 PERSONAL DOCUMENTS
Personal diaries and letters – Provide first-hand accounts of people’s thoughts,
feelings, and experiences.
Interviews and participant observation – Offer detailed insights into behaviour and
social interactions.
Historical tapes and transcripts from interviews – Give access to past events and
perspectives.
Commercially published autobiographical sources – Provide life stories and
reflections, but may be influenced by publication purposes (e.g., bias or image).
📄 EVALUATING THE QUALITY OF PERSONAL
DOCUMENTS
Authenticity – Is the document genuine, and is the author ( ?واقعیConfirm the origin
and authorship.)
Credibility – Is the information accurate and free from distortion?
(Consider factual accuracy and whether it reflects the writer’s true feelings.)
Representativeness – Is the document typical or biased?
(Some documents, like letters, may not survive, leading to selective survival and
possible bias.)
Meaning – Is the content clear and understandable?
(Watch for damaged documents, abbreviations, or unclear language.)
📄 PUBLIC DOCUMENTS
Public documents – Generally considered authentic and meaningful as they are
officially produced.
Credibility – May be biased, depending on the purpose or organisation that created
them.
Representativeness – Can be difficult to determine, as not all documents are preserved
or accessible.
📄 ORGANISATIONAL DOCUMENTS (Expanded but
Concise)
Organisational documents – Generally seen as authentic and meaningful, as they are
produced by organisations.
Credibility – May be biased, especially in documents like annual reports, mission
statements, press releases, and advertisements, which aim to create a positive image.
Authenticity & Representativeness – Can be challenging, particularly with online
documents, where origin and completeness may be uncertain
📄 MEDIA OUTPUTS
Media outputs – Authenticity can sometimes be difficult to verify, especially with
mass-media sources.
Representativeness – Usually not a major issue, as it is often possible to identify the
larger population (e.g., newspapers or magazines).
Meaning – Content is generally clear and easy to understand, making it useful for
analysis.
Visual documents in business research
Visual documents (images) can be used in both quantitative and qualitative research.
Photographs may be analysed as part of content analysis or used in interviews and experiments.
They provide useful context and deeper insight.
Three common uses are:
1. Illustrative – used to support and make findings more engaging and easier to understand.
2. As data – treated as research data and analysed to identify patterns or themes.
3. As prompts – used to encourage discussion and more detailed responses from
participants
Documents as ‘texts’
As authors use various means (such as language, tone, and structure) to influence their readers,
researchers who use documents to understand an organization should support their analysis with
other data sources to improve validity and reduce bias. Documents may reflect particular
viewpoints or hidden agendas rather than objective reality.
Interpreting documents can be done using methods such as qualitative content analysis,
semiotics, and historical analysis. These approaches help researchers examine meanings,
symbols, context, and underlying messages within the text, leading to a deeper understanding of
the document’s purpose and significance.
Referring to websites
As academic researchers pay increasing attention to both the content and changes in data
collected online or related to websites, the date the website was accessed should always be
included in the reference. This is important because online content can be updated, moved, or
removed over time.
When researchers follow up on your research or review your findings, they may find that the
websites are no longer available or that the content has changed. Including the access date
improves transparency, credibility, and allows others to better verify your sources.
QUANTITATIVE CONTENT ANALYSIS
A method used to quantify content:
‘a research technique for the objective, systematic and quantitative description of the manifest
content of communication.’ It focuses on analysing visible (manifest) content rather than hidden
meanings.
Objectivity – aims to reduce researcher bias by using clear rules and procedures.
Being systematic – follows a structured process, ensuring consistency in how data is
collected and analysed.
Content analysis can be used to analyse research ethics codes of nine social scientific
associations at four different levels:
Ethical tone – the overall attitude or approach to ethics
Ethical values – the key values promoted (e.g., honesty, integrity)
Specific ethical principles – formal guidelines or rules
Specific research practices – how these principles are applied in practice
Research questions for content analysis
Content analysis focuses on examining patterns in communication by asking key questions such
as:
Who – who is being reported or represented (e.g., individuals, groups, organisations).
What – what issues, topics, or events are being reported.
Where – where the issue is reported (e.g., types of media or platforms).
Location – where the coverage appears within the content (e.g., headlines, main articles,
or smaller sections).
How much – the amount or frequency of coverage given to a topic.
Why – the reasons or purpose behind why the issue is reported, including possible
agendas or motivations.
What is computer-aided content analysis?
While manual content analysis requires researchers (coders) to carefully read through text and
identify key terms or themes themselves, computer-aided content analysis uses specialized
software to automatically search, identify, and count words, phrases, or patterns within large
bodies of text.
This method is especially useful when dealing with large datasets such as newspaper articles,
academic journals, online content, or social media posts. Because most modern texts are
available in digital form, researchers can analyse data more quickly and efficiently without
relying entirely on manual coding.
Computer-aided content analysis improves:
Speed – large volumes of data can be analysed in a short time
Consistency – reduces human error and bias in coding
Efficiency – allows researchers to focus more on interpreting results rather than
collecting data
However, it may still require human input to interpret meaning, context, or nuanced language
(like sarcasm or tone), which software might miss.
Common software used includes:
[Link] – [Link]
NVivo – [Link]
home
QDA Miner Lite – [Link]
software/freeware/
WordStat – [Link]
📘 Coding
Coding is the process of organising and categorising data in content analysis by assigning labels
or codes to specific words, phrases, or sections of text. It helps researchers identify patterns,
themes, and relationships within the data. Coding must be systematic, use clearly defined
categories, and ensure that categories are mutually exclusive and exhaustive to maintain
accuracy and consistency.
📘 Coding Schedule
A coding schedule is a structured tool or document used to record the codes assigned during
content analysis. It contains predefined categories and allows researchers to systematically
capture data, often using counts or checklists. The coding schedule ensures that data is organised
and consistent, making it easier to analyse and compare results.
📘 Coding Manual
A coding manual is a detailed set of instructions that guides researchers on how to apply codes
during content analysis. It includes definitions of categories, rules for coding, and examples to
ensure clarity and consistency. The coding manual helps reduce bias and ensures that different
coders interpret and classify data in the same way.
📊 Content Analysis Measures – Expanded Notes
🔹 What content analysis reveals
Content analysis helps a researcher understand what is important and how meaning is
communicated in a text.
1. What the text establishes as relevant
o Identifies key topics, themes, or issues highlighted in the text
o Shows what the author focuses on most
o Helps determine the main purpose of the document
2. The priorities portrayed through the text
o Reveals what is emphasized vs what is ignored
o Indicates the importance placed on certain ideas
o Helps understand the agenda or focus of the author
3. The values conveyed in the text or image
o Shows beliefs, attitudes, or viewpoints
o Can reflect bias, ideology, or cultural perspectives
o Helps identify whether the message is positive, negative, or neutral
4. How ideas are related
o Shows connections between concepts
o Helps identify patterns, themes, or arguments
o Reveals structure and flow of information
🔹 What content analysis measures
This focuses on how the content is examined in a systematic and objective way:
1. What is contained
o Looks at specific words, phrases, or ideas
o Identifies the presence of certain themes or concepts
o Example: counting how often “conflict” appears in a text
2. How frequently it occurs & order of occurrence
o Measures how often something appears (frequency)
o Looks at sequence or positioning (what comes first, repeated emphasis)
o Helps show importance and emphasis
3. Positive or negative views
o Identifies tone or sentiment (favourable vs unfavourable)
o Useful for analyzing opinions, attitudes, or bias
o Example: whether employees are described positively or negatively
4. Proximity of ideas (logical association)
o Examines how close ideas appear together in the text
o Helps identify relationships or connections between concepts
o Example: linking “stress” with “workplace conflict”
📊 Potential Pitfalls in Devising Coding Schemes – Expanded
Notes
When developing a coding scheme in content analysis, researchers must avoid common mistakes
that can affect the accuracy and reliability of results.
🔹 Discrete dimensions (no overlap)
Each dimension (category or variable) must be clearly separate
There should be no conceptual or empirical overlap between dimensions
Overlapping dimensions can confuse coders and lead to inconsistent results
👉 In simple terms: Each category should measure only one unique idea
🔹 Mutually exclusive categories
Categories within each dimension must not overlap
A single piece of data should fit into only one category
If categories overlap, coders may not know where to place the data
👉 In simple terms: One item = one category only
🔹 Exhaustive categories
All possible responses or data must be covered
Coders should always find a suitable category for every piece of data
Missing categories can lead to incomplete or biased analysis
👉 In simple terms: No data should be left uncategorized
🔹 Clear instructions to coders
Coders must be given detailed guidelines
Instructions should explain:
o What each dimension means
o How to interpret the content
o What to consider when assigning codes
Lack of clarity leads to inconsistency and subjectivity
👉 In simple terms: Everyone coding should understand and apply rules the same way
🔹 Clear unit of analysis
The unit of analysis must be clearly defined
Researchers must distinguish between:
o The media item (e.g., newspaper article)
o The event or topic discussed in the item
Confusion here can lead to incorrect coding
👉 In simple terms: Be clear about what exactly you are analyzing
📊 Quantitative Content Analysis – Expanded Notes
Quantitative content analysis is a method used to systematically and objectively analyze
content by quantifying it (e.g., counting words, themes, or patterns).
✅ Advantages
🔹 Transparent and objective method of analysis
The process is clear, structured, and systematic
Reduces researcher bias because it follows specific coding rules
Results can be replicated and verified by other researchers
👉 In simple terms: It is fair and consistent
🔹 Used for longitudinal analysis
Can analyze data over long periods of time
Helps identify trends and changes in content
Useful for studying developments (e.g., changes in media reporting)
👉 In simple terms: Good for tracking changes over time
🔹 An unobtrusive method
Does not involve direct interaction with participants
Data is usually already available (e.g., documents, media)
Avoids influencing the subject being studied
👉 In simple terms: Does not disturb or affect people
🔹 Highly flexible method
Can be applied to many types of data (texts, images, media)
Suitable for different research topics and contexts
Can be adapted for various research purposes
👉 In simple terms: Can be used in many different ways
❌ Disadvantages
🔹 Issues of authenticity, credibility, and representativeness
Data sources may not always be reliable or accurate
Some documents may be biased or incomplete
May not represent the full population
👉 In simple terms: Data may not always be trustworthy or complete
🔹 Coding manuals cannot avoid interpretation
Even structured coding requires some level of judgment
Different coders may interpret content differently
Can affect consistency of results
👉 In simple terms: It is not completely objective
🔹 Risk of invalid inferences
Researchers may draw incorrect conclusions from the data
Counting words/themes does not always reflect true meaning
👉 In simple terms: Numbers don’t always tell the full story
🔹 Not suitable for answering “Why” questions
Focuses on what is present, not why it occurs
Lacks deep understanding of motives or reasons
👉 In simple terms: It explains what, not why
🔹 Atheoretical
May lack connection to theory
Focuses more on description than explanation
Can limit deeper academic insight
👉 In simple terms: May not explain the bigger picture
🔹 Can separate meaning from context
Breaking content into categories may remove it from its original context
Important nuances or meanings may be lost
👉 In simple terms: Meaning can be lost when content is simplified
📊 Checklist for Evaluating Documents – Expanded Notes
This checklist is used in content analysis to assess the quality, credibility, and usefulness of a
document before using it as data.
🔹 Who produced the document?
Identify the author or organization responsible
Consider their background, expertise, and credibility
Helps determine if the source is reliable and trustworthy
👉 In simple terms: Know who created the information
🔹 Why was the document produced?
Understand the purpose of the document
Was it created to inform, persuade, entertain, or report?
Purpose can influence the content and tone
👉 In simple terms: Ask what is the goal of this document
🔹 Authoritative position
Determine if the author has knowledge or expertise on the topic
Experts or professionals are more likely to provide accurate information
Non-experts may provide less reliable content
👉 In simple terms: Is the author qualified to speak on this topic?
🔹 Is the material genuine?
Check the authenticity of the document
Ensure it is not fake, altered, or misleading
Confirm the document’s origin is valid and legitimate
👉 In simple terms: Is it real and trustworthy?
🔹 Hidden agenda (bias or slant)
Determine if the author has a specific interest or bias
Look for signs of persuasion, manipulation, or one-sided views
Identify any “slant” in the way information is presented
👉 In simple terms: Is the author trying to influence the reader?
🔹 Representativeness (typicality)
Assess whether the document is typical of its kind
If it is unusual, determine how and why it differs
Helps avoid drawing conclusions from unrepresentative data
👉 In simple terms: Is this document normal or an exception?
🔹 Clarity of meaning
Ensure the content is clear and understandable
Look out for vague language, missing information, or ambiguity
Clear meaning is necessary for accurate analysis
👉 In simple terms: Can you easily understand it?
🔹 Corroboration
Check if the information can be verified using other sources
Compare with other documents or evidence
Helps confirm accuracy and reliability
👉 In simple terms: Can this information be confirmed elsewhere?
🔹 Alternative interpretations
Consider whether the document can be understood in different ways
Be aware of other possible explanations or perspectives
Justify why you accept one interpretation over others
👉 In simple terms: Could someone interpret this differently?
🧠 Easy way to remember
When evaluating a document, ask:
Who? → author
Why? → purpose
Can I trust it? → authenticity & credibility
Is it biased? → hidden agenda
Is it typical? → representativeness
Is it clear? → meaning
Is it confirmed? → corroboration
Are there other views? → interpretations