CONTENT ANALYSIS- Krippendorff
Chapter 1: History of Content Analysis
1.1 Origins & Precursors
Content analysis — the systematic reading of texts, images, and symbolic matter — traces back
to 17th-century theological inquiries, when the Church sought to control the spread of
non-religious printed material. Though the term itself did not appear in English until 1941, its
practice is far older.
The first well-documented quantitative analysis of printed matter occurred in 18th-century
Sweden, centred on the Songs of Zion controversy. Scholars debated whether the hymns
harboured dangerous or dissenting ideas. The dispute generated foundational methodological
questions still relevant today: Should meanings be read literally or metaphorically? Do symbols
carry different meanings in different contexts? This was an early rehearsal of core
content-analysis debates.
1.2 Quantitative Newspaper Analysis
Early 20th-century journalism schools demanded empirical inquiry into the press. Researchers
measured column inches devoted to subject-matter categories as a way of claiming 'scientific
objectivity.' They believed quantified facts were irrefutable. This approach — quantitative
newspaper analysis — attempted to 'reveal the truth about newspapers.'
Limitations: It was largely journalism-driven, relied on crude subject-matter categories, and
made inferences without acknowledging the analyst's own conceptual contributions to the data.
1.3 Early Content Analysis (1930s–1940s)
Four contextual factors drove content analysis into a new phase:
● Post-1929 economic crisis → concern that mass media caused social breakdown, yellow
journalism, rising crime
● Rise of new electronic media (radio, later TV) — could not be treated simply as
extensions of print
● Fascism and political challenges to democracy were linked to the power of radio
propaganda
● Emergence of the behavioural and social sciences with new empirical methods
Key conceptual advances in this period:
● The psychological concept of "attitude" added evaluative dimensions (pro/con,
favourable/unfavourable) to analysis — opening the door to the systematic assessment
of bias.
● Interest in social stereotypes (Lippmann) brought questions of media representation into
focus — how groups such as minorities or nations were depicted.
● Symbol analysis: Harold Lasswell, drawing on psychoanalytical theory of politics,
classified symbols into categories (self/others, indulgence/deprivation) and developed a
World Attention Survey comparing symbol use across elite newspapers of different
nations.
● Berelson & Lazarsfeld codified the field in their 1948 text and the 1952 book Content
Analysis in Communications Research — the first systematic methodological treatment.
Transition from quantitative newspaper analysis to content analysis involved:
● Eminent social scientists asking new kinds of questions
● Theoretically motivated, operationally defined concepts (symbols, values, propaganda
devices replaced crude subject categories)
● New statistical tools borrowed from survey research and experimental psychology
● Content analysis data becoming part of larger multi-method research efforts
1.4 Propaganda Analysis (World War II)
WWII forced the most important and large-scale application of content analysis. Rather than
identifying 'propagandists' (as pre-war researchers had), analysts now needed military and
political intelligence. Two analytical traditions emerged from this period:
● Lasswell's group: focused on newspapers and wire services from abroad; worked on
sampling, measurement, reliability and validity of content categories.
● The FCC group: analysed domestic enemy (Axis) broadcasts to understand and predict
events inside Nazi Germany
Key lessons learned from WWII propaganda analysis:
● Content is NOT inherent in communications. People read texts differently. The sender's
intent may have little to do with how audiences receive the message. Context, individual
needs, preferred discourses, and social situation all shape meaning.
● Content analysts must infer phenomena they cannot directly observe — this is the
primary motivation for using the technique. The question is not 'what is in the text?' but
'what can be legitimately inferred from available texts?'
● Analysts need elaborate models of the systems in which communications occur —
earlier analysts had treated messages as inherently meaningful unit-by-unit; propaganda
analysts succeeded only when they viewed messages in the full context of the diverse
people using them.
● Quantitative indicators are often too shallow for political intelligence. Qualitative analysis
can be equally systematic, reliable, and valid. George's post-war validation work
confirmed this.
The key conceptual shift: From treating content as 'shared' or 'manifest' (Berelson's later
term), the FCC analysts moved toward understanding the motivations of specific communicators
and the interests they served. The notion of 'preparatory propaganda' — preparing domestic
populations for planned military actions — became a powerful analytical key.
1.5 Generalisation Across Disciplines (Post-WWII)
After WWII, and following Berelson's 1952 codification, content analysis spread widely:
● Mass communication: Lasswell's World Attention Survey; Gerbner's 'cultural indicators'
project — analysing violence profiles in TV fiction across nearly two decades, tracing
how groups (women, minorities, elderly) were portrayed.
● Psychology (four areas): (1) inferring motivational/personality traits from verbal records;
(2) open-ended interview and focus group data; (3) interaction process analysis of
communication (Bales); (4) semantic differential scales (Osgood) for cross-cultural
comparison of meaning.
● Anthropology: myths, folktales, riddles; component analysis of kinship terminology.
Methods converge with those of content analysts.
● History: systematic analysis of historical documents; inferring past events from surviving
texts.
● Literary studies: identifying authors of unsigned documents; author attribution problems.
A conference noted two convergences across disciplines: (a) a shift from analysing 'content' to
drawing inferences about the antecedent conditions of communication; (b) a shift from
measuring volumes of subject matter to counting symbol frequencies, then to relying on
contingencies (co-occurrences).
1.6 Computer Text Analysis
Late 1950s onwards: computer languages suitable for literal (non-numerical) data processing
opened new possibilities. Key developments:
● General Inquirer system: dictionary-based computer content analysis, applied across
political science, advertising, psychotherapy, and literary analysis. Most influential early
system.
● TextPack: a more general, multilingual successor to General Inquirer.
● Sedelow proposed using a thesaurus rather than a dictionary, as it better reflects
'society's collective associative memory.'
● WordNet (George Miller): a computer-traceable network charting word meanings.
● NVivo and [Link]: interactive-hermeneutic software for computer-aided qualitative
analysis — human coders translate text into categories; computers handle volume and
summarising.
● DeWeese bypassed costly transcription by feeding newspaper typesetting tapes directly
into a computer — a landmark step.
Core debate: Computers are powerful at scanning large volumes of text reliably, but their
operations remain confined to their programmers' conceptions. Without human intelligence,
computer analysis cannot point to anything outside of what it processes. Computers have no
environment of their own — they operate in users' worlds without understanding them. The most
productive model: computers as aids, not replacements, for human reading and inference.
1.7 Qualitative Approaches
Krippendorff questions the quantitative/qualitative distinction — all reading of texts is ultimately
qualitative, even when results are later converted to numbers. Nevertheless, qualitative
approaches offer alternative systematic protocols.
Key qualitative approaches:
● Discourse analysis: Focuses on text above the sentence level. Examines how
particular phenomena (racism, ideology, stereotypes) are represented.
● Social constructivist analysis: Focuses on how reality is constituted through human
interaction and language — e.g., how facts are constructed, how emotions are
conceptualised
● Rhetorical analysis: Examines how messages are delivered and with what effects.
Identifies structural elements, tropes, styles of argumentation, speech acts.
● Ethnographic content analysis: Does not avoid quantification but encourages
categories to emerge from reading. Focuses on situations, settings, styles, images,
nuances; is emic (insider perspective) rather than etic.
● Conversation analysis: Starts from recorded verbal interactions in natural settings;
analyses transcripts as records of 'conversational moves' toward collaborative
construction of conversation.
Shared characteristics of qualitative/interpretive approaches:
● Require close reading of relatively small amounts of textual matter
● Involve re-articulation (interpretation) of texts into new analytical narratives accepted
within particular scholarly communities
● Analysts acknowledge working within hermeneutic circles — their own socially/culturally
conditioned understandings constitutively participate (Krippendorff calls these
interactive-hermeneutic approaches)
Chapter 2: Conceptual Foundation
2.1 Definition of Content Analysis
Content analysis is a research technique for making replicable and valid inferences
from texts (or other meaningful matter) to the contexts of their use.
Unpacking the definition:
● Technique: Involves specialised, learnable procedures independent of the researcher's
personal authority. It is a scientific tool.
● Replicable: The most important form of reliability. Researchers working at different
times and under different circumstances should get the same results applying the same
technique to the same data. Procedures must be explicitly stated and applied equally to
all units.
● Valid: Research findings must be open to scrutiny and upheld against independently
available evidence. Validity is testable; objectivity is not (hence Krippendorff drops
Berelson's requirement of 'objectivity').
● Inferences to contexts: Results do not simply describe the text itself but point to
phenomena beyond the text — conditions, meanings, uses, effects.
Three Types of Definitions in the Literature
● Type 1 — Content inherent in text (Berelson): Content is 'manifest' in the message,
waiting to be described. Berelson's definition: 'objective, systematic and quantitative
description of the manifest content of communication.' Problem: implies content is
'contained' inside messages; treats meaning as a single, discoverable thing; uses the
container metaphor.
● Type 2 — Content as property of the source (Holsti; Osgood): Content analysis tied
to inferences about the states or properties of the source. Holsti's encoding/decoding
paradigm: describes 'what, how, to whom' to infer 'who, why, with what effects.'
Limitation: putting sources in charge of validity may not capture all communicators'
intents; ignores the analyst's own conceptual contributions.
● Type 3 — Content emerging in researcher-context relationship (Krippendorff):
Content emerges in the process of a researcher analysing a text relative to a particular
context. The analyst's conceptual participation is acknowledged. This is Krippendorff's
preferred definition.
Critique of Berelson: His requirement that content be 'manifest' effectively limits analysis to
what everyone agrees on — excluding expert interpretation (e.g., a psychiatrist reading a
patient's story differently). The 'container metaphor' for meaning is misleading and still abounds
in communication research (Krippendorff). It implies one meaning per message and makes
findings immune to invalidating evidence.
Gerbner extends Berelson by arguing that mass-media messages carry the imprint of their
industrial producers and that audiences are affected by statistical properties of messages they
are not consciously aware of — which privileges the content analyst's reading over audience
readings.
2.2 Six Key Epistemological Features of Texts
Krippendorff argues that the popular measurement model for content analysis (borrowed from
mechanical engineering) is misleading. Instead, six features of texts must be recognised:
● (1) Texts have no objective, reader-independent qualities: A text does not exist
without a reader. Messages do not exist without an interpreter. Data do not exist without
an observer. Meanings are always brought to texts by someone. There is nothing
inherent in a text.
● (2) Texts do not have single meanings: Texts can be read from numerous
perspectives — counted, categorised, analysed for metaphors, given psychiatric,
political, or poetic interpretations. All may be valid but different. Believing in one 'correct'
meaning is a naive entailment of the container metaphor.
● (3) Meanings need not be shared: Demanding 'common ground' would restrict content
analysis to trivial manifest aspects. Psychiatrists, anthropologists, conversation analysts
legitimately read texts in ways that differ from how the subjects/authors read them.
Content analysis is only in trouble when expert readings fail to acknowledge how
designated audiences actually use the texts.
● (4) Meanings speak to something outside the text: Communications inform, invoke
feelings, cause behavioural changes. Texts refer to events elsewhere, objects that no
longer exist, ideas in people's minds. The analyst must look outside the physicality of the
text. This is also the key limitation of computer analysis: computers operate in users'
worlds without understanding those contexts.
● (5) Texts have meanings relative to particular contexts: Once an analyst chooses a
context, the diversity of interpretations may be reduced to a manageable number.
Different disciplines (psychiatry, political science, literary studies) construct different
contexts for the same text. Context determines what questions are answerable.
● (6) Content analysis demands specific inferences: Analysts must draw specific
inferences from texts to their chosen context. Systematic reading narrows the range of
possible inferences about unobserved facts, intentions, mental states, and effects.
2.3 The Conceptual Framework
Krippendorff's framework has three purposes: (1) prescriptive — guides research design; (2)
analytical — facilitates critical comparison of published analyses; (3) methodological — points to
performance criteria. The framework has six components:
● Body of text (data): The starting point. Unlike experimental data, content analysis data
are not generated to answer a specific research question — they are texts meant to be
read by others. Text results from reading and re-articulation.
● Research question: The target of inference. Must be answerable from texts
(abductively inferable), must concern currently inaccessible phenomena, must allow for
at least in-principle validation. Not the same as a statistical hypothesis — content
analysis answers questions about extratextual phenomena.
● Context: The analyst's constructed world in which texts make sense. Embraces all
knowledge — theories, propositions, empirical evidence, intuitions — that the analyst
applies. Context must be made explicit; without it, any reading would be equally justified.
Different disciplines construct different contexts.
● Analytical construct: Operationalises what the analyst knows about the context. Takes
the form of 'if-then' rules of inference guiding the analyst from texts to answers.
Functions like testable mini-theories of a context. Must model the context of use — not
just any statistical technique applied to text.
● Inferences: The central accomplishment of content analysis. The kind of inference is
abductive (not deductive or inductive) — proceeding across logically distinct domains,
from text features to answers about extratextual phenomena. Analogous to Sherlock
Holmes's reasoning.
● Validating evidence: Any content analysis must be validatable in principle (even if
infeasible in practice). Without this requirement, analysts could pursue questions that
yield results backed only by their own authority.
Types of Inference
● Deductive: From generalisations to particulars. Logically conclusive. Not central to
content analysis.
● Inductive: From particulars to generalisations. Statistical generalisation. Not central to
content analysis.
● Abductive (key type): Across logically distinct domains — from text features to
extratextual phenomena. Probabilistic, not conclusive, but strengthened by taking
contributing conditions into account. This is the logic of Sherlock Holmes and of content
analysis. Toulmin's argumentation theory: the analytical construct (the warrant) plus
reliable application justifies the inference; the context backs the warrant.
Validating Evidence — Key Points
● Validation may be infeasible (e.g., wartime intelligence, detecting lying politicians) or
impossible (e.g., inferences about deceased authors, historical events)
● The requirement of in-principle validatability prevents self-serving or purely abstract
categorisations
● Ex post facto validation (George's validation of FCC wartime inferences using captured
Nazi documents) is valuable — it advances the technique for future use
● Validation requires repeated use of the same categories and constructs across studies
— ad hoc designs contribute little to the discipline
● Berelson and Lazarsfeld's principle: there is no point in counting unless the frequencies
lead to inferences about surrounding conditions. Raw counts without abductive inference
are not content analysis.
2.4 Contrasts with Other Research Methods
Content analysis has four distinguishing features compared with other social science methods:
● Unobtrusive / non-reactive: Content analysis avoids the contamination introduced
when subjects know they are being studied (contrast with surveys, experiments,
interviews). Texts are produced independently of the analyst's interest. This also
prevents sources from strategically misleading the analyst (e.g., if Goebbels had known
how his broadcasts were being decoded, he would have adapted).
● Handles unstructured data: Unlike surveys or experiments that impose structure and
suppress individual variation, content analysis works with naturally generated texts in
diverse formats. The chief advantage: it preserves the conceptions of the data's sources.
● Context sensitive: Content analysis processes data that are meaningful, significant,
and representational to others. Unlike context-insensitive methods (surveys,
experiments) that strip data from their original context, content analysis proceeds by
reference to contexts of its own — making inferences more likely to be relevant to actual
users of the texts.
● Handles large volumes of data: Explicit, repeatable procedures allow many coders (or
computers) to process quantities of text far beyond what any single individual can
analyse — from thousands of newspaper editorials to decades of television
programming (Gerbner). Today, electronic full-text databases make volume virtually
unlimited, shifting the bottleneck from data access to theory, methodology, and software.