0% found this document useful (0 votes)
5 views174 pages

Introductory Research Methodology Guide

Uploaded by

ama.tomcsanyi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views174 pages

Introductory Research Methodology Guide

Uploaded by

ama.tomcsanyi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

INTRODUCTORY RESEARCH

METHODOLOGY

Jozsef Zoltan Malik

BUDAPEST METROPOLITAN UNIVERSITY

Budapest, 2019

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


I. Introduction: The Scientific Way of Thinking

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Why to bother with RM?
Some Key Questions in Research Methodology

However, Research Methodology is useful not only for PhD students and
scholars but for everyone who wants to make a professional work on a subject.
If you are familiar with academic standards, your way of thinking and your
performance may be more professional. Why?
An academic essay, thesis, or research project should demonstrate that its
author(s) has
 a critical, perceptive and constructive analysis of the subject;
 skills in the gathering and analysis of information and report presentation;
 surveyed literature relevant to the topic;
 carried out original and significant work in a field.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Scientific Orientation
”Most people learn about the “scientific method” rather than about the scientific
attitude. While the “scientific method” is an ideal construct, the scientific
attitude is the way people have of looking at the world. Doing science includes
many methods; what makes them scientific is their acceptance by the scientific
collective.”
Fredrick Grinnel: The Scientific Attitude (1987)

It is more important to grasp the


orientation or the attitude of science
rather than to follow a certain “scientific
method.”
 The scientific method is not one thing; it is
a collection of ideas, rules, techniques,
and approaches used by the scientific
community.
 The scientific orientation tends simulta-
neously to
 be precise and logical,
 adopt a long-term view,
 Be flexible and open ended, and
 be willing to share information widely.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Science and Technology
Science is an intellectual activity carried on by
humans that is designed
- to discover information about the world and
society in which humans live, and
- to discover the ways in which this information
can be organized into meaningful patterns.
”Scientific knowledge”
 The systematic observation of natural/social events and conditions in order to
discover facts about them and to formulate principles, laws, and mechanisms
based on these facts.
 The organized body of knowledge that is derived from such observations and that
can be verified or tested by further investigation.
 Any specific branch of this general body of knowledge, such as physics, biology,
geology, or economics, sociology, and political science.

Technology is the process by which humans modify


nature to meet their needs and wants.
”Technological knowledge”
 "...the know-how and creative processes that may assist people to utilise tools, resources
and systems to solve problems and to enhance control over the natural and made
environment in an endeavour to improve the human condition." (UNESCO, 1985).
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Make a Difference: Science and Technology
In our ages science and technology often appear together (S&T)
referring to advanced technology based on new scientific principles,
the two are different in many aspects.

Science Technology
Subject unchangeable changeable
Goal knowing the ”general” knowing some ”concrete”
How it works? inside: ”facts from facts” outside: ”patented invention”
Activity Theoria: end in itself Poisesis: end in sthg else
Method abstraction modelling concrete
Process conceptualization optimization
Innovation from discovery invention
Type of Result law-like statements rule-like statements
Time Perspective long-term short-term

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


What Does It Mean to Be Scientific? (#1)
It’s difficult to define what is meant ”science”. The simplest and very
brief definition would be this one:
An attempt to identify and test empirical generalizations.

”Empirical:”
 This term refers to the facts or events of the real world: that which exists
and can be known through experiences;
 Empirical statements refer to what is or is not true and can be confirmed
or disproved by experience.
”Empirical” vs. ”Normative:”
 Much of what we might believe about things is not empirical, but rather
normative – it reflects our judgments about what should be.
 Normative questions deal with value judgments, that is, questions of what
is good or bad, desirable or undesirable, beautiful or ugly.
Reformulating Normative Questions as Empirical:
 Though scientific research can deal with normative phenomena, but it
can do so only indirectly as it seeks to answer empirical questions;
 This can be done by taking the normative questions that motivate our
interest and reformulating them as empirical questions.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


What Does It Mean to Be Scientific? (#2)
It’s difficult to define what is meant ”science”. The simplest and very
brief definition would be this one:
An attempt to identify and test empirical generalizations.

”Generalization:”
1. Scientists seek to make statements about entire classes of objects, not
just individual cases, though the observation must be of individuals
 Example:
1. Mr. Smith has only a grade school education and does not vote;
2. Ms. Jones has an advanced degree and always votes.
?
 To make a generalization that people with more education are more
likely to vote than people with less education, we need to collect
information on a large number of people from many places and
across time.
2. The main purpose of science is to explain and predict, and thus
scientific explanation requires generalizations.
 The generalizations made in the social sciences are almost never
absolute.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


How to Reformulate?
There are two techniques to use:
To change the frame of reference:
 The gist is moving from a normative judgment to a question about the
normative judgements some person or persons make.
To ask empirical questions about the assumptions behind
normative judgements.
 Most normative judgments are based at least in part on beliefs about what
is empirically true. But are these assumptions true?
 This way of reformulating is more sophisticated than the first one, it
requires much deeper analysis.
Examples:
1. Would it be a good idea to legalize drugs? (Normative)
 Do most Germans/Dutch favour legalization of drugs? (Frame)
 Would legalization of drugs decrease the occurrence of other crimes? (Ass.)
 How much would legalization of drugs increase the frequency of addiction? (Ass.)
2. The United States should continue to send troops to the third world to
attempt to restore order. (Normative)
 Nations in the European Union favour the U.S. sending of troops in most cases.
(Frame)
 The support of peacekeeping activities with U.S. troops generally has not resulted in
long-term prevention of disorder in the past. (Ass.)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


On Generalization

DEDUCTION INDUCTION
Hypothesis:
 This is an educated guess
based upon observation.
Theory  A hypothesis is an explanation Observation
for a phenomenon which can
be tested in some way which
ideally either proves or dis-
proves the hypothesis.
Hypothesis  The goal of the researcher is to
Patterns
rigorously test the terms of the
hypothesis.

Theory:
 The explanation or a model for
Observation a phenomenon. Hypothesis
 A conceptual framework that
explains existing observations
and predicts new ones.
 Theories not only describe why
Conformation or how the phenomenon oc- Theory
curred but also guide the way
for further research.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


On Generalization (con.)
Two directions of theorizing:

DEDUCTION
Logical Inferences

1. Modus Ponens
2. Generalization 1. If A, then B. 1. AB
2. A is the case. 2. A is true
3. Therefore, B is true. 3. So B is true.

INDUCTION
Karl Popper: Modus Tollens:
1. All scientific knowledge is hypothetic 1. If A, then B.
 Not can be generalized 2. B is not the case.
3. Therefore, A is false.
2. Modus Tollens, and not MP
 Falsification 1. AB
2. B is false.
3. So A is false.

Can never be absolutely certain Conclusion: Even the most well-established and popular
one has observed ALL instances. scientific theory can never be proved – it can only be disproved.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Making Science: some characteristics

Science is empirical. Science relies on experience


more than authority, common sense, or logic.
Science is objective. Objectivity means that same
conclusion should be arrived if same observation is
made.
Science is self-correcting. Because science is
empirical, new evidences may contradict the old ones.
Science is progressive. Because science is empirical
and self-correcting, it is also progressive.
Science is tentative. Science never claims to have the
whole truth. New information may make current
knowledge obsolete.
Science is parsimonious. Use the simplest
explanation to account for a phenomenon.
Science is concerned with theory. Develop theory of
how something works.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


About ”Progressive” Science
The Evolution of Science:
Traditional view: Linear
and cumulative (follows a
direct path from past to
present, adding at each
point to the achievements
of earlier generations).
Kuhn’s view: Scientific
development is not
smooth and linear;
instead it is episodic—that
is, different kinds of
science occur at different
times.
 The most significant
episodes in the development
of a science are normal
science and revolutionary Paradigm Shift is when a
science. It is also cyclical significant change or scientific
with these episodes re- revolution happens.
peating themselves.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Discipline
Discipline: A particular branch of scientific knowledge.
A discipline has seven basic characteristics:
1. Focus of study
2. Paradigm
3. Reference disciplines
4. Principles and practices
5. Research agenda
6. Education
7. Professionalism
The emergence of a new discipline:
 Kuhn: When enough significant anomalies have occurred against a current paradigm, the
scientific discipline is thrown into a state of crisis.
 During this crisis, new ideas, perhaps ones previously discarded, are tried.
 Eventually a new paradigm is formed, which gains its own new followers, and an
intellectual "battle" takes place between the followers of the new paradigm and the hold-
outs of the old paradigm.
 The new paradigm may lead to a new discipline.
Conclusion: We have to take a post-positivist attitude in science.
 Science does not itself aim at some grand goal such as the Truth; rather individual scientists seek to solve
the puzzles they happen to be faced with.
 There is no logic of science or fixed scientific method. Instead scientists make discoveries thanks to their
training with exemplary solutions to past puzzles.
 Nor is it cumulative, since revolutionary science typically discards some of the achievements of earlier
scientists.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Make a Difference: Science and Ideology
In daily life, we encounter many doctrines and ideologies that share features with
social theory.
 Both tell us why things are the way they are: why crime occurs, why some
people are poor but not others, why divorce rates are high in some places, etc.
 Both contain assumptions about the fundamental nature of human beings
and of the social world.
 Both offer systems of ideas or concepts, and both interconnect the ideas.
Ideology Science (Theory)
Certainty of answers absolute, certain answers with few tentative, conditional answers that
questions are incomplete and open ended
Type of knowledge closed, fixed belief system open, expanding belief system
Type of assumptions implicit assumptions based on explicit, changing assumptions
faith, moral belief, or social position based on open, informed debate
and rational discussion
Use of normative merger of descriptive claims, separation of descriptive claims,
statements explanations, and normative explanations, and normative
statements statements
Empirical evidence selective use of evidence, resistance Seeking repeated tests of claims,
to contrary evidence changing based on new evidence
Logical consistency contradictions and logical fallacies highest levels of consistency,
avoiding logical fallacies
The idea that simple is better;
Occam’s Razor or everything else being equal, a
Conspiracy Theory
Parsimony: social theory that explains more
with less complexity is better.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Theory vs. Ideology: Examples
Divorce:

Science (Theory): Ideology


– Family are strongest when they – Society is facing a moral
have resources (income, decay leading to divorce,
education, housing, maturity, women working outside the
respect etc.) and low stress home, and loss of the
(constant employment, happy “traditional family.”
marriage, good health, etc).

The debate over Evolutionism and Creationism:

Evolutionism Creationism
– Evolution qualifies as a scientific – Creationism relies on sacred teachings or
theory because of its logical writings that believers accept as being
coherence, openness, integration absolute truth and largely do not question.
with other scientific knowledge,
and empirical tests.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


II. On Research

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


What is Research?
Definition:
Research is the systematic process of collecting and analyzing
information (data) in order to discover new knowledge or expand
and verify the existing one.

Research Characteristics:
 Controlled – in exploring causality in relation to two variables, the study
must be set in a way to minimise the effects of other factors affecting the
relationship.
 Rigorous – be scrupulous in ensuring that the procedures followed to find
answers to questions are relevant, appropriate and justified.
 Systematic – the procedures adopted to undertake an investigation follow a
certain logical sequence ... Different steps cannot be taken in a hazardous
way.
 Valid and verifiable – whatever is concluded on the basis of the findings
must be correct and can be verified by the researcher and others.
 Empirical – any conclusions drawn are based upon hard evidence gathered
from information collected from real-life experiences or observations.
 Critical – critical scrutiny of the procedures used and the methods
employed.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


What is Research? (#2)
The Reasons for revisiting theories:
 Geographical: May have been tested only in one country/region
 Example: 1) Behaviour patterns of urban residents – are they
replicated in rural areas? 2) Theory established using US data could
be tested in another country.
 Social: May have been established on the basis of the experience of one
social group only
 Example: Theory based on men’s experience – does it apply to women?
 Temporal: May be out of date
 Example: Theory on youth culture established in the 1980s – is it still
valid?
 Contextual: May have been established in fields other than we are working.
 Example: 1) Foucault’s theories on power are based on studies in a
hospital – are they relevant in the tourism industry?
2) Dahrendorf’s class-conflict theory can be applied to firms?
 Methodological: May have been tested using only one methodology.
 Example: Conclusions from a qualitative study could be tested
quantitatively

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


L&T as interdisciplinary Study (#1)
Leisure & Tourism studies a special thematic framework:

The linkages between people, organisations and


services/facilities/attractions consist of
processes such as:
 Link A – market research and political activity;
 Link B – marketing, buying, selling, employing,
visiting/using services;
 Link C – planning and investment.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


L&T as interdisciplinary Study (#2)
Disciplines vary in terms of their primary focus of attention within this
system:
 Psychology and social psychology are focussed
primarily on the people element, with some concerns
with links A and B.
 Political science is concerned mainly with
organisations and with link A to the people;
 History can cover the whole system – but much of
historical research in leisure
studies has also had the same focus as political
science;
 Economics at the macro-level is concerned with the
whole system, while micro-economics is located around
Link B, where the market process is at work;
 Sociology is concerned primarily with the people
and with Link A and with  Link A – market research
organisations; and political activity;
 Geography’s basis is the interaction between the  Link B – marketing,
human parts of the system and the environment; buying, selling, employing,
 Applied disciplines, such as planning, management visiting/using services;
and marketing, are based in organisations, then move
 Link C – planning and
along links A and C to the other elements of the
system; investment.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research: From Idea to Theory/Paper

Purposes of Research:
- EXPLORATORY
- DESCRIPTIVE
- EXPLANATORY

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research Types And Their Purposes
Exploratory Research: e.g. field experiment
What you want to know What you will do
Become familiar with the basic facts, Create a general mental picture of conditions;
setting, and concerns.  Formulate and focus questions for future
research;

Results:
Generate new ideas, conjectures, or hypotheses;
Determine the feasibility of conducting research;

Descriptive Research: e.g. census of population


What you want to know What you will do
The characteristics of the observed Provide a detailed and accurate picture;
phenomena or of how it works or behaves. Locate new data that contradict past data;
Create a set of categories or classify types;
Clarify a sequence of steps or stages;
Results:
Document a causal process or mechanism;
Report on the background or context of a situation.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research Types And Their Purposes (cont.)
Explanatory Research:
What you want to know What you will do
The causes of the observed Explain what factors or conditions are
phenomena/events that have occurred or are causally connected to a known outcome.
happening.
Results:
 Predictive:
The future outcome of current conditions Predict what outcome will occur as a result of
or trends. a set of known factors or conditions.
 Prescriptive:
The things that can be done to bring about Prescribe what should be done to prevent sg
some outcome from happening or to bring sg about.
 Normative:
What is best, just, right, or preferable, and
Adjudicate among different understandings
what ought to be done (or not done)
of how sg should be or what ought to be
done, by considering the arguments of others
and rational reasons of one’s own.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research Question
Definition: Any scientific research begins with a question that
the research is intended to answer.

Which attributes make potentially good RQ?


Clarity: The RQ must be specific enough to give direction to the
research, and general enough that suggests what a possible answer
would be.
Testability: The RQ must be one that can be potentially answered by
empirical inquiry.
 It must be an empirical question, not a normative question;
 Whether the necessary investigation can be devised and carried
out with the resources available.
Theoretical significance: Answering the RQ should potentially
increase our general knowledge and understanding of the topic.
Practical Relevance: Answering the RQ should be useful in some
real-life application.
Originality: This does not mean that a research question must he
completely new, but it does mean that the answer should not be so
well established that there is little reason to expect a different
outcome.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research Question
RQ is the most important step in a research
 Research project should pose a question that is relevant in the real world.
 The topic should be consequential for political, social, or economic life;
 However, no more ”valuable” or more ”fruitful” research topic than another one:
Research question often comes from a currant social problem (e.g. drug consumption)
as well as from a theoretical dilemma (“What we have now is not quite right/good
enough – we can do better ...”)
Research project should make a specific contribution to an identifiable
scholarly literature by increasing our collective ability to construct verified
scientific explanations of some aspects of the world.
 Research question is
 the starting point of the research process that initiates and generates the action of
research, and
 the end-point of the feedback.
 Research process may be drifted in two ways:
- empirical-inductive way: OBSERVATION  CORRELATION(S)  GENERALISATION
- analytic-deductive way: HYPOTHESIS  ANALYSIS  EXPLANATION(S)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Types of Research Questions
Descriptive
”What do we mean by…”? Exploratory or
 What’s going on? Where, with whom? Who are concerned? Descriptive
 What is the context for it? Research
 Whose accounts/perspectives do we have?

Explanatory
 What helps to explain what happens? What factors highlight/underlie it?
 What concepts and theories help to understand it?
 What causes it?
Evaluative (explanatory with prediction)
 How useful/effective is it? Explanatory
 Does it work? (Over what period of time?) Research
Strategic (explanatory with prescription)
 What are the implications?
 Should it affect on practice or policy priorities? How and Why?

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research Question: Examples
Example#1: At a school parents complain much about a respected teacher,
and they want her to be dismissed. The headmaster is in trouble because
the teacher has been working for the school for ten years.
It’s worth making an exploratory research.
 The basic question: What’s wrong with the teacher?
 Are the grades of the pupils are getting wrong?
 Are the teacher too strict?  field experiment
 Is there something problem in the relationship between the
teacher and the class?
 Is there any problem in the personal life of the teacher?
 How to improve class work/curriculum?

Example#2: Research on Terrorism


Descriptive questions:
 What is terrorism?
 What type of individual becomes a terrorist?
 How are terrorist actions similar to or different from
military actions?
Explanatory questions:
 Why do people or groups engage in terrorism?
 How do the media affect our view of terrorism?
Explanatory questions with prediction/prescription:
 Where are future terrorist attacks likely to occur?
 How can governments control terrorism effectively?

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Fallacies in framing research questions
RQ must be
 sufficiently narrow or specific to permit empirical investigation;
 formulated as a type of question that asks what you’re really interested in answering.

However, a RQ that satisfies these requirements may still not be researchable.


The most common errors when a question is framed so that
1. it begs another question; 4. it is metaphysical;
2. it presents a false dichotomy; 5. it is a tautology.
3. it asks about fictional event;

Examples:
 Why was American slavery the most awful that was ever known?
 Though it is an explanatory question, it contains another question: Whether
American slavery was the most awful that was ever known? (1.)
 It also contains a false dichotomy (2.): you need to choose between two answers that
are neither mutually exclusive nor collectively exhaustive.
Would US president Franklin D. Roosevelt have decided to drop A-bombs on
Japan had he still been in office in August 1945?  Fictional question (3.)
Was the WW2 inevitable if the Parisian Peace Pact in 1919 had been more
respectful for the Germans?  Metaphysical question (4.)
Was Mr. X an unsuccessful president because he was moving against the tide of
history?
 We know nothing about the tide of history  Mr. X was unsuccessful, because he was
unsuccessful (this is the prejudice of the researcher)  Tautology (5.).
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
III. Literature Review

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Why Literature Review? (#1)
Suppose you want to be familiar with a topic
 For a while, when you are getting in, you’ll be confused, surely.

What you will meet:


 Diversity of opinions,
 Different stances and approaches,
 Agreements & Disagreements,
 Diversity of terminology (especially in
new areas or different disciplines), ...
 Partial relation to your work, etc.
Your job as a researcher:
Build a conceptual framework (on your mind first)

Rule of Thumb: Your Used ideas, results, ... from others


work won’t be accepted must be properly referenced
(for publication) without Facilitate contextualization
a proper study of and Ethical issue – Plagiarism, Reputation
comparison with related
works.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Why Literature Review? (#2)
What is the place of LR?

 Bring clarity and focus to your research problem


Helps you understanding the subject
Helps you to conceptualize your research problem
Helps identifying relationships with existing body of
knowledge
 Improve your method
How the others have approached the problem
Which methods others have used and faced difficulties
Broaden your knowledge base in your research area
You need to know where we are and where the gaps are
Help identifying trends
It is convenient to know what are the hot research
topics in the area
what are the assessment criteria in use
Contextualize your findings
How your results fit into the existing body of knowledge
How do your results differ from others

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Sources
Traditional Sources:
 Books
 Journal papers
 Conference papers
 Reports.

Online Sources:
 Most publishers are making their
products accessible online (subject to
subscription)
 Reference databases are also available
online
 Some scientific associations give
online access to their publications for
subscribers/members
 There is a trend at universities to
subscribe packages guaranteeing
access to contents from multiple
publishers.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Exploration of Sources (#1)
Primary source of information:
 In the study of liberal arts and humanities,
primary source always serves as an original
source of information about the topic.
 It includes any artefact, document, diary,
manuscript, autobiography, recording, and other
source of information that was created at the time
under study.
Secondary sources of information:
 Secondary sources serve us to have a comprehensive view about the
subject of the research without reviewing each document exhaustively.
 Two types of secondary sources:
 Analytic secondary documents aim at exploring and representing just
one specific work; they involve book review, biography, documentary,
recension, etc.
 Synthetic secondary documents involve all the sources related to the set
of a topic or field.
Factorgraphical data is direct information to our question from
encyclopaedias, dictionaries, or Google.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Exploration of Sources (#2)
Access to papers available via the
web :

 Google Scholar

[Link]

 JSTOR

[Link]

 and many mores…

[Link]

[Link]

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Source Criticism (#1)
When making a literature survey … pay special attention
to the reliability of the sources:
 Is it coming from a prestigious journal?
 Was it presented in a serious peer-reviewed
conference?
 Are there other related references?

In a more general context, there are no good or bad sources.


 Whether a source is relevant or not depends on what you would like to get to
know from it.
 Example: Suppose your research interest is how witch burning took place in
the 16th century in Europe.
 An eyewitness account by a farmer’s wife who saw a witch being burned
would be a good source.
 However, her account would perhaps not be a good source if you want to
know how the witch defended herself in front of the judge in the town.
 The town mayor who hated witches would be a bad source, indeed.
 Court records can tell you much more about the statements the witch
made in the court.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Source Criticism (#2)
Some questions to put to the source. To examine your source
critically, you should answer the following questions regardless of
whether they are ancient or from yesterday.
1. What kind of source is it?
 What type of source is it? Is it a letter, a diary, a law, a treaty, etc? The
type of source can explain why it contains the information it does.
2. Who wrote the source? And Why?
 Who and why did the author(s) produce the text? What could be the
motive for writing the narrative?
 Always take into consideration whether the creator of the source could
have a special interest in lying, exaggerating, or bending the truth.
3. When is the source from?
 Look at the event and circumstances under which the source were
created.
 If it was written a long time after the event took place, there is a risk
that the creator of the source has forgotten what actually happened or
that he or she does not remember correctly.
4. Is it a primary or secondary source?
 It is important to know this, because a secondary source is always just
a copy or an interpretation of the original source.
 That’s why it is a good idea to go to the primary source if it is
accessible.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Source Criticism (#3)
5. Is the information first-hand or second-hand?
 Was the creator present when the event happened or
was he or she told about it by someone else?
 If the creator was present and saw what took place, we call this a first-
hand account. If he or she heard about it from others, we call the source a
second-hand account.
 First-hand sources are better, but we should still remain critical because
even if the creator was an eyewitness to an event, he or she could still
have an interest in exaggerating, lying or not telling the whole truth. Or he
or she could have forgotten the details.
6. Who is the source addressed to?
 Pay attention to the recipient of the source. The creators of the sources
may have had an interest in writing as they did because they knew who
would read it.
7. Is the source backed by other sources?
 One source is not enough to explain events. After all, the information in
the source may not be correct. There are often several different accounts
that tell about the same event.
 In addition, the accounts often disagree about what actually happened.
Make use of your toolbox for source criticism to assess which sources
come closest to the truth.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Annotated Bibliography
Annotated Bibliography is not Literature Review
 An annotated bibliography is a list of citations to books, articles, and
documents. Each citation is followed by a brief (usually about 150 words)
descriptive and evaluative paragraph, the annotation. The purpose of the
annotation is to inform the reader of the relevance, accuracy, and quality
of the sources cited.
 Examples:

 Unlike an annotated bibliography, which provides information on one source at a time, a LR


offers a generalized picture of what scholars have thought and written about your topic.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
How to make a Literature Review? (#1)
PRIMARY GOALS:
 A LR provides a scholarly context for the argument
researchers propose and support in their paper.
 Most often (e.g. thesis, book), a literature review is
formatted to appear as a separate section of paper,
preceding the body. It helps readers perceive how your
argument fits into past and present scholarly discussion
of your subject.
STEPS IN PREPARING A LITERATURE REVIEW:
1. Survey and determine the important representative samples of the scholarly
literature on the topic.
2. Summarize the contents of those works, taking notes on
 the author’s field of expertise;
 the types of evidence the author relies on (e.g., case studies, narratives, statistics,
primary sources) and the reliability of this evidence;
 the author’s stance and approach;
 the author’s arguments (indicating which are most convincing and which are less so)
 the author’s contributions to scholarly discussion of the topic
[Link] you are sufficiently familiar with the individual works you have
examined, look for patterns among them. Determine how they compare and
contrast.
(Hint: constructing a chart to organize your findings visually often makes it easier to
discern various kinds and degrees of similarity and difference.)
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
How to make a Literature Review? (#2)
ORGANIZING LITERATURE REVIEW:
No one standard format. Nor specific instructions about the length
of a LR (a general rule of thumb is that it should be proportionate to
the length of your entire paper).
The main parts of LR in general (recommendation):

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


IV. Conceptualisation

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Conceptualisation (#1)
Science starts and ends with theories, and all theories have some main parts:

 Assumptions: An untested starting point or belief in a theory that is


necessary in order to build a theoretical explanation.
 All theories contain built-in assumptions, and thus they are at least in part
subjective.
 Theories:
Subjective (Assumption)  Intention to Objectivity (Research Characterisation)

 Subjectivity: an integral part of your way of thinking that is conditioned by your


educational background, discipline, philosophy, experience and skills.
 Bias: a deliberate attempt to either conceal or highlight something.
 Post-Positivist Philosophy: Science itself is definitely not ”value-neutral.”

 Example: To study racial discrimination and prejudice,


we might assume that
 people have it in varying degrees, and some people
may not have it at all;
 some persons discriminate people in other racial
groups but not the ones belonging to their own
racial group;
 racial prejudice persists over time in a person and
does not instantly appear or disappear.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Conceptualisation (#2)
The Evolution of Conceptual Framework:

Operationalisation involves
moving from the abstract What are the main
to the specific. features/components
to grasp?  Different
sources of conceptualisation

Sketch a concept map 


Example: Negotiations

Concept Pyramid:
-Simple/complex concepts
- Concept Clusters
- Typology
- Ideal Type

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Conceptualisation (#3)
Different Sources of Topics and Conceptualisation:
 Personal Interest
- The researcher may be personally involved in
an activity
- he may be a member of a particular social group

 Popular/Media
- Common beliefs - Internet
- Popular issues - Audiovisual materials (e.g. TED)

 Literature Review
- To contextualise the subject
- To reach conceptual framework

 Reports, Datasets
- To gather relevant information
- To recognise patterns, trends, agendas, etc.

 Other Disciplines
- To have different stances
- Alternative approaches

 Brainstroming, Conferences
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Conceptualisation (#4)
Concept Map:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Conceptualisation (#5)

 Concepts: The keywords and notions occur in RQ and


hypotheses need definitions.
 Definition: Theoretic concept is an idea that is thought
through, carefully defined, and made explicit in a theory.
 Simple vs. Complex Concepts:
 Simple concepts have only one dimension and vary along a
single continuum.
- Problem: the same concept can different in different
theories/disciplines: ”value” really means different in
economics and sociology.
 Complex concepts have multiple dimensions or many
subparts.
 Example: ”Democracy”
 Regular, free elections with universal suffrage;
Each dimension
 An elected legislative body that controls government; varies by degree
 Freedom of expression and association.
 Concept Cluster: A collection of interrelated concepts that share
common assumptions, refer to one another, and operate together in
a social theory.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Conceptualisation (#6)
 Typology: A theoretical classification that is created by cross-classifying or
combining two or more simple concepts to form a set of interrelated subtypes.
 Example: ”Family”
”A family (from Latin: familia) is a group of people related either by
consanguinity (by recognized birth), affinity (by marriage or other
relationship), or co-residence (as implied by the etymology of the
English word "family") or some combination of these.”
 Ideal Type:
 It is a broader, more abstract concept that organises a set of
more concrete concepts;
 It outlines the central aspects of what is of interest;
 It is a pure, abstract model that tries to define the core of the
phenomenon in question.
 Example:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Operationalisation
Theory is a set of empirical generalizations about a topic, but its statements are
too general to test directly and the investigated relationship between its abstract
concepts are complex and not directly observable.  Task: we must break it
down to more specific term.
This is done by testing hypotheses (educated guesses based upon observation).
In research process hypotheses actually are statements about variables.
Indicators are phenomena which point to the existence of the concepts.
Variables are components of the indicators which can be measured. So,
variable is an empirical property that can take on two or more different values.
The Unit of Analysis in the hypothesis – the objects that the hypothesis
describes.
Each variable in a hypothesis must have an operational definition: a set of
directions as to how variable is to be observed and measured.
Example:
THEORY: Socioeconomic status affects political
participation.
HYPOTHESIS: The higher a person’s income,
the more likely he or she is to vote.
OPERATIONAL: The higher a respondent’s
answer when he or she is asked in a survey,
"What is your household’s annual income," the
more likely that person will answer "Yes" when
asked, "Did you vote in the election last week?"

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


An Example: from concept map to operationalisation (#1)

Concept map:

Draft A Draft B

Draft C

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


An Example: from concept map to operationalisation (#2)

Definition and Operationalisation:


Concept Definition Operationalism
Participant:
in Leisure Person who engages in Participation in activity identified as
relatively freely chosen ‘leisure’ at least once in preceding
activity during leisure time. year
in Tourism Person who travels away from Travel for leisure purposes at least
home for leisure purposes 40km from home with at least one
overnight stay in preceding 3
months.
in Politics Person who join a group of Size matters – counting participants
people for declaring and
achieving a common goal

Indicators:
- The total of leisure/tourism/political - The utility of costs and satisfaction
-Price experience
- Access - Affordable price (L&T) - The range of facilities and its ”benchmark”
- The quality of organisation & costs
the reaction of police force - Reports & survey
(politics)
- Individual - Age - Age last birthday
- Income - Annual household income
attributes

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Hypotheses (#1)
Causation – Diagrams of Causal Relationships:
a. Positive relationship:

Positive relationship: An
association between two
concepts or measures so
b. Positive & Negative relationships: that as one increases, the
other also increases, or
when one is present, the
other is also present.

Negative relationship: An
association between two
concepts or measures so
c. Complex relationship: that as one increases, the
other decreases, or when
one is present, the other is
absent.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Hypotheses (#2)
Types of hypotheses and examples:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Hypotheses (#3)
Types of variables and examples:
 Definition:
Independent variables are those
Dependent variables are the presumed in the theory underlying
effects or consequences the hypothesis to be the cause.

Although this distinction is sometimes difficult to make, in most hypotheses it


is apparent; explicitly expressed in the statement, e.g. ”causes”, ”leads to”, etc.
 Control (Intervening) variables are additional variables that might affect the
relationship between the independent and dependent variables, however their direct
effects are always excluded.
 When control variables are used, the intent is to ensure that it is not these
variables that are in fact responsible for the variations observed in the
dependent variable.
 Examples:
HYPOTHESIS: Urban areas have HYPOTHESIS: Education and political
lower crime rates than rural areas. participation are positively related.
INDEPENDENT VARIABLE: Urbanization INDEPENDENT VARIABLE: Education
DEPENDENT VARIABLE: Crime rates DEPENDENT VARIABLE: Political participation
CONTROL VARIABLE: Region/State CONTROL VARIABLE: Age
UNIT OF ANALYSIS: Geographic areas UNIT OF ANALYSIS: Individuals

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Hypotheses (#4)
Further examples (Try to fill out the cards):

HYPOTHESIS: People from ethnic group A are more likely to


commit crimes than people from ethnic group B.
INDEPENDENT VARIABLE:
DEPENDENT VARIABLE:
CONTROL VARIABLE:
UNIT OF ANALYSIS:

HYPOTHESIS: A person who has no job is more likely to suffer


from mental illness than the one who has.
INDEPENDENT VARIABLE:
DEPENDENT VARIABLE:
CONTROL VARIABLE:
UNIT OF ANALYSIS:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Hypotheses (#4)
Further examples (A possible solution):

HYPOTHESIS: People from ethnic group A are more likely to


commit crimes than people from ethnic group B.
INDEPENDENT VARIABLE: Ethnicity
DEPENDENT VARIABLE: Crime rate
CONTROL VARIABLE: Age
UNIT OF ANALYSIS: Individual

HYPOTHESIS: A person who has no job is more likely to suffer


from mental illness than the one who has.
INDEPENDENT VARIABLE: Employment
DEPENDENT VARIABLE: Health
CONTROL VARIABLE: Addiction
UNIT OF ANALYSIS: Individuals

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Statistical Hypotheses
A common feature of the statistical method is the concept of the null hypothesis,
referred to by the symbol H0.
 It is based on the idea of setting up two mutually incompatible hypotheses, so
that only one can be true.
 Example: Either more people play tennis than golf or the number of people
who play tennis is less than or equal to the number who play golf – if one
proposition is true then the other is untrue.
The null hypothesis usually proposes that there is no difference between two observed
values or that there is no relationship between variables. There are therefore two
possibilities:
H0 – Null hypothesis: there is no
significant difference or relationship
H1 – Alternative hypothesis: there
is a significant difference or
relationship.
Note! Usually it is the alternative hypothesis,
H1, that the researcher is interested in, but
statistical theory explores the implications of
the Null hypothesis.

In terms of Research Methodology, this is very much a deductive approach: the


hypothesis is set up in advance of the analysis, possibly within a theoretical
framework.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
V. The ”Methodology” of Data

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


From Concepts to Measurement
To systematically investigate ideas and conjectures about some aspect of political
reality, it is necessary to move from abstract world into the empirical world.
 This means thinking about what data or evidence is relevant to answering your
RQ.

There are at least three main steps to take:


1. Conceptualisation:
 Aim: To have a conceptual definition of what it is that we want to investigate.
 Validity: Is our definition of the concept valid? Is it intuitive enough to adopt?
 Example: What it is that we mean by the term ”democracy”?
2. Operationalisation:
 Aim: To construct indicators we can use to tap into our definition.
 Validity: Are they valid indicators of our conceptual definition?
 Example: We are giving two characteristics of democracy (competition &
participation), our task is to find the corresponding indicators to observe.
3. Measurement:
 Aim: To find data sources to measure our indicators.
 Reliability: Are the indicators reliable?
 Example: To use the Freedom House Index.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Example: The Term ”Democracy” (#1)
Conceptual Definition: Let’s adopt a minimalist definition of democracy:
Definition: ”Democracy is a system of electing governments through
competitive elections with full adult franchise.”
 Competition and Participation
Validity: Competition and participation are necessary but not sufficient
criteria to label a political regime as democracy.
 We may acknowledge that there is more to democracy than simply
holding elections, but that it is also true that democracy cannot exist
without elections.
Operationalisation: This involves moving from the abstract to the specific.
Competition is the difference in vote share between the winning party and
the second placed party in the same election.
 If one party wins by a huge, we can say the election is not very
competitive; otherwise, it is.
Participation is the proportion of the adult population that turns out to
vote.
 If a country some groups are not eligible to vote (e.g., women, different
ethnic or racial or religious groups), or there are barriers to
registration, or where the turnout is very low, this could be examples
of restricted participation.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Example: The Term ”Democracy” (#2)

Operationalisation: This involves moving from the abstract to the specific.


Validity:
 First, is the margin of victory in an election a valid indicator of
competition?
 And is the level of adult participation a valid indicator of
participation?
 We may object to democracy being defined in such narrow terms,
but concede that the indicators are valid.
Measurement:
 To measure competition, we may decide to inspect official election
returns that are published after all the votes have been counted.
 To measure participation, we may count the number of votes that have
been cast in the election and compare this to the most recent census
estimates of adult population size in the country.
 The key issue is to do with reliability.
 Are the election results an accurate reflection of what actually
happened in the election?
 May be they have been manipulated in some way (Electoral fraud).
 Does the country have reliable population estimates?
 Are there existing, up-to-date sources that we can use?

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Example: The Term ”Democracy” (#3)
In Practice – The Methodology of Freedom House Index.
FH is a US-based NGO that conducts research and advocacy on
democracy, political freedom, and human rights.
They use a checklist with two main categories and several
characterisations (operational definition):
1. Political Rights
A. Electoral Process
B. Political Pluralism and Participation
C. Functioning of Government
2. Civil Liberties
D. Freedom of Expression and Belief
E. Associational and Organizational Rights
F. Rule of Law
G. Personal Autonomy and Individual
Rights

Each characteristic point from A to


G is assessed by scores on an
interval scale, and finally these
scores add up to a total score.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Supplamentary: Checklist for Freedom House Index

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Validity
Whatever type of data is collected or used,
the researcher will have to confront issues
of validity and reliability.

The Types of Validity:


 Face validity simply means that, on the
face of it, the indicator intuitively seems like
a good measure of concept.
 Content validity examines the extent to which the indicator covers the full
range of the concept, each of its different aspects.
 Construct validity examines how well the measure conforms to our
theoretical expectations by examining the extent to which it is associated
with other theoretically relevant factors.
 Example:
1. Face validation  competition and participation are really important
elements of democracy
2. Content validation  we might argue that this minimal definition of
democracy lacks content validity, we should give some additional
indicators (e.g. freedom of press).
3. Construct validity  how strongly
the missing indictor does change
our theoretic expectations?
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Reliablility
Even if we have a valid indicator, if we are
not capable of measuring it accurately, we
may end up with unreliable data.
In brief: Data can be valid but not
reliable, and can be reliable but not
valid.
Example: Corruption
[Link] definition: What is
meant by ”corruption”? How to
grasp it?
[Link] definition: Focus on
the perception of corruption, not on
the incident  Usually this may be
observable.
Dilemma: Perception  Reality?
(Does the perception that a country
is corrupt mean that it really is
corrupt?)
[Link]: what data could provide
a better source of information?
(Transparency International regu-
larly carries out a survey which asks
businesspersons to rate the level of
corruption in different countries.)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Coding of Data (#1)
The scale (levels) of measurement: A system for organizing information in the
measurement of variables
Scale Presumption Operations on data
Nominal Nominal scale identifies property, it The change of symbols in the set
contains categories of a variables of data.
about ”what kind” not ”how much.”
Ordinal Ordinal scale identifies a difference All kinds of change is permitted
among categories of a variable and that hold up the order of data
allows the categories to be rank
ordered as well.
Interval Interval scale identifies differences The determination of distances
among variable attributes, ranks and differences
categories, and measures distance
between categories.
Ratio (or Ratio scale is to determine the The determination of ratios and
similitude of ratios with using an multiples
Absolute) absolute zero.

Remark: In most practical situations, the distinction between interval and ratio
levels makes little difference.
 The point is that on interval scale we use solely arbitrary zeros, the zeros are
only to help keep score. On absolute scale zero is absolute, and this makes it
possible to state relationships in terms of proportion or ratios (i.e. we can
measure continuous data).
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
The Coding of Data (#2)
Examples:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Coding of Data: Supplamentary
Other Examples: On ”absolute zero”:

Zero degree in Fahrenheit or Celsius is


not the absence of any heat but is just
a placeholder to make counting easier.

If there were a true zero, the actual relation among temperature numbers would be a ratio.
For example, 25° to 50° Fahrenheit would be “twice as warm,” but this is not true because a
ratio relationship does not exist without a true zero. We can see this in the ratio of boiling to
freezing water temperatures. The ratio is 5.625 times higher in Fahrenheit, 100 times in
Celsius, and 0.366 times in Kelvin. The Kelvin scale has an absolute zero (K is a ratio scale,
whilst F and C are just interval scales), and its ratio corresponds to physical conditions.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Beyond our study….
There are several other issues that are worth investigating related
to data. But they require a much deeper and more advanced
studies about the methodology of data.
Selection Issues:
How much information is enough to set up an ”optimal”
dataset for a research?
Internal Validity:
 Concerned with whether a relationship between two
variables is causal or not: “If we assume that A causes B,
how can we be sure that it is A that is responsible for a
variation in B and not something else that produces an
apparent causality?”
External Validity:
 “Can the results be generalised beyond the specific context in which they
were produced?” For instance, if you perform a case study on the most
successful company, will your results hold for other companies, as well?

Big Data: Data Mining & Machine Learning

Kenneth Cukier: Big data is better data

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


VI. Research Design

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research Design in a Broader Sense

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Logical Method of Research
However, in our presentation now, research design
refers to the building blocks by which we propose
explanations, and the main forms as explanations
appear in social researches.

The scientific life can be seen as a never-ending chain of


Research Question  Paradigms  Hypotheses 
 Research Study  Theory  New Research Questions
In this never-ending adventure deduction and induction rotate
each other. From existing theories we deduce hypotheses, and new
theories are induced by the results of researches.
In Deduction: Researcher has some great theories that are ready to be
applied. Though these theories are rivals and the interpretations lead to
competitive results, but the results are considered as those with general
scope.
In Induction: Researcher makes attempt to customise theories, he/she
sets up the setting of explanations according as to what the nature of
the subject of the research is; how it works; how it is endured or
altered.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Three Reasearch Traditions in Socials
Research Traditions/Strategies in Socials:
 Any question concerning the subject of Socials is from three directions:
 Micro-level: Rational Choice approach
 Macro-level: Structural (Holistic) approach
 Mezo-level: Middle-Range Theory – A constructive approach

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Quantitative vs. Qualitative Research (#1)
The quantitative paradigm of researches:
 The results of social researches can be expressed in
numbers, they appear in ”hard data”.
 The basis for analysis is statistics from which we can
infer laws with general scope, and so we can make
predictions.
 The quantitative researchers collect data, which are
existed out there independently of the researchers, like
miners who dredge up minerals.
The qualitative paradigm of researches:
 The basis for analysis is an interpretive approach with the
purpose not to discover general laws of socials but to
explore and to understand social phenomena and to create
theories rested on the inductive way of generalization.
 The results of social researches cannot be expressed in
”hard data” but in ”soft data” involving stories, case studies,
words rather than numerical data.
 The qualitative researchers generate data, and similarly to
travellers, they are talking, asking, discovering, and finally
interpret their observations. That’s why the qualitative
researches are not independent of their researches.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Quantitative vs. Qualitative Research (#2)
The framework of quantitative researches:
 The goal of quantitative researches is to present and to
verify laws and tendencies by testing hypotheses set up
in advance.
 Thus the framework of the research is always structured
and static, usually is done under artificial
circumstances.
 The size of the sample is large.

The framework of qualitative researches:


 The framework of qualitative research is non-structured
and flexible, the hypotheses are generated during the
process of the research.
 The goal of the research is to scrutinise the cases
investigated under the circumstances where they
are/were appeared.
 The size of the sample is small, based mainly upon
observations, case studies and interviews.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Nomotethic and Idiographic Explanations
Quantitative researches use nomothetic analyses
Goal: The extention of findings based on a sample.

The criteria for establishing causation in nomothetic analyses:


(1) The variables must be empirically associated, or
correlated,
(2) the causal variable must occur earlier in time than the
variable it is said to affect, and
(3) the observed effect cannot be explained as the effect of a
different variable.

Qualitative researches use idiographic analyses


Goal: An effort to understand the features of some
individual, organic unit (person, organisation, culture)

(1) The idiographic model aims at a complete under-standing


of a particular phenomenon, using all relevant causal
factors,
(2) these explanations must be logical and plausible, and
(3) should be pointed out that the other alternative
explanations were not the case.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Quantitative vs. Qualitative Research Design

Quantitative Reserach Qualitative Research


Nomothetic: Idiographic:
Concepts The extention of findings based on An effort to understand the features of
a sample. some individual, organic unit
(person, organisation, culture)
Predetermined and operationalised Constructed during the research
1) Anonymous (transformed into 1) Names with capitals
variables) 2) Tend to select paradigmatic cases
Cases 2) Tend to select randomly 3) Systematic process analysis
3) Cases are independent from
each other
1) Increase N whenever possible 1) Keep N low (usually N < 10)
(N > 30 might as well be okay) 2) Increase number of variables in
Variables 2) Reduce the number of order to make the description
variables in order to avoid thicker (full accounts)
undetermined research design
Typical Statistical Method Case studies,
Surveys Interviews
Methods

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Forms of Explanation
By ”forms of explanation” we mean causal relations embodied in claims in
the explanation. The main forms of explanations:
 Laws: If the causal relation
 is not particular but typical (repeated),
 the cases can be comparable, and
 is based upon nomothetic analyses.
 Mechanisms are causal patterns, models, which make connections between
observable phenomena or events.
 The scope of mechanisms is not universal like that of laws, and
 mechanisms are rested upon deductive-analytic models or idiographic
analyses.
 Trends or Tendencies: Mechanisms that, beyond the causal relations of
phenomena, want to present a regularity about a set of phenomena or the
society itself. Though, we need to be very carefully about tendencies, we may
accept them if
 the period of the subject of the research is not very long,
 the phenomena in the set are associated and investigated enough, and
 the size of the observed cases is large enough.
 Examples: 1) Business cycles, 2) Malthusian Thesis (1798): The unchecked
population growth is exponential while the growth of the food supply was
expected to be arithmetical  - disease, starvation, and war, and – a need
of population control.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
The Time Dimension in Social Research
 Cross-sectional Research:
Any research that examines
information on many cases at
one point in time.

 Time-Series or Longitudional
Research:
Any research that examines
information from many units or
cases across more than one
point in time.

 Panel Study:
A time-series research in which
information is about the
identical cases or people in each
of several time periods.

 Cohort Study:
A time-series research that
traces information about a
category of cases or people who
shared a common experience at
one time period across
subsequent time periods.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Research Design and Scientific Outlooks

Positivist Post-positive Interpre(ta)tive Humanistic/Postmodern

Reason for To discover laws To discover and To understand To express the subjective
so people and describe
research can predict and understand meaningful self, to be playful, and to
control events social action social action entertain and stimulate
Social Objective, Objective and Objective &
reality subjective as Subjective
(naive) realism critical realism intrinsically linked
Yes but not easy It is not separate No, focus on human
Knowability Yes from human
to capture subjectivity subjectivity
Relationship Dualist: Knowledge is The aim is to No objective knowledge is
btw. scholars influenced by understand
inductive possible
and subject the scholar subjective
(deductive) knowledge
Good is based upon is set against is embedded in the has aesthetic properties
evidence theoretic
precise considerations context of fluid and resonates with
observation (by contrast) social interactions people’s inner feelings
Forms of Laws Plausible Contextual Empathic
knowledge inferences
(causal relations) knowledge knowledge
(mechanisms)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Types of Research Design (#1)
Experimental Designs:
 Field Experiment: Cheap, flexible way of observing phenomena directly
 Lab or Simulation: A combination of artificial and empirical observations
of phenomena  The findings will be qualified as ”soft data.”

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Types of Research Design (#2)
Unobtrusive Designs:
 The researcher is an observer.
 To describe and explain social phenomena.
 The description is often complemented, through some indicators, by a middle-
range analysis.

Descriptive Designs:
 Provide a detailed and accurate picture about the characteristics of the
observed phenomena or of how it works or behaves.
 Report on the background or context of a situation.
Historical Designs:
 They can help us to discover the influence of certain events on other in two
ways:
 Cross-sectional: by contextualising events so as to enable us to better
understanding;
 Longituduional: by helping us to understand how the timing and sequence of
social actions and process impact on their nature and on the development of
subsequent actions and events.

Case Study: Focusing on a single case, the case can be intensively examined.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Types of Research Design (#2)
Deductive-Analytical Design:
 A standard secondary research method based on the processing of the
literature and the interpretation of secondary (published) data.
 Complex theoretical considerations as the basis for deduction.
 Deductive theorizing with the aim of explaining social phenomena.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Types of Research Design (#3)
Comparative Designs:
 They are perhaps most widely used research design in science.
 Within the comparative framework based upon empirical data it is common to
distinguish to main methods of research designs.

Empirical-Quantitative Method (Research Methods: t-test, ANOVA)


 It rests on large-N studies (i.e. the size of the sample is large).
 Its results can be expressed in numbers, they appear in ”hard data”.
 it uses nomotethic analyses.
Two main forms: - To study recorded statistical data (e.g. GDP, salary,
population, etc.); - To make surveys.
Empirical-Qualitative Method (Research Methods: MSSD, MDSD)
 It rests on small-N studies (and tending to select paradigmatic cases).
 The results of social researches cannot be expressed in ”hard data”
but in ”soft data” involving stories, case studies, words rather than
numerical data.
 It uses idiographic analyses.
 Two main forms: - To make case studies, - To make interviews.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


An Example
Methodological assessment:
• Declared Aim: To tend to a positive
research design in the spirit of
Dankwart Rustow
 The authors mentions but avoid
normative aspects.
• Research Designs:
 Descriptive
 Historical
 Comparative using Empirical-
qualitative Method:
 Small-N
 Cross-sectional: more
regions, but at the same
time (”the third wave of
democratization”)
 Based upon country case
studies.

Munch-Leff: Modes of Transition and Democratization: South America


and Eastern Europe in Comparative Perspective, 1997

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Types of Research Design (#4)
The Correlation Designs:
 They require only collecting data on an independent and a dependent variable
and determining whether there is a pattern of relationship.
 It is usually advisable also to collect data on other potentially relevant variables
and statistically control for them.
 Two types of statistical tests are used in this research design:
 Chi-Square test that can be used to compare variables if one of them is
categorical (nominal or ordinal).
 Correlation and regression analysis that can be used to compare interval
variables.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


VII. Relevant Research Methods in Socials

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


A Simplified Typology

Research Methods

Secondary Comparative Primary Research Methods


Research Methods Methods

Qualitative Research Quantitative Research

E.g.: A Content Analysis may belongs to the scope of each methods in practice.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Unobtrusive methods

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Methods for Historical Researches
Historical Research:
It involves the investigations of events that
 either have an important impact on subsequent
developments,
 or provide an opportunity of testing the implications of a
general theory.
 Assumption: Social phenomena are considered as a result of
history (Genealogical Postulate);
 The challenge of this research design:
 Precise description, and to find and point out paradigmatic
cases.
 The chronologic issue: to mindmap the causations in course
of events and to find the crucial cause(s): to seek out a
”Point of No Return,” the moment from which there is an
immanent logic of the events up to the denouement.
 Work out an appropriate Typology or Ideal Type by which
you make a synthesis.
Three main methods for Historical Researches:
1. Historical Events Study: We have one case and one period of time.
2. Historical Process Research: We have one case but many time periods.
3. Comparative Historical Research: We have two or more social settings or
groups (e.g. countries) and our aim is to infer something from the comparison.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Examples: Historical Researches
Examples:
 Historical Events Study:
 This method is excellent to explore how the socially
shared representations of past events creates and
maintains social identity and contributes to the
intergroup dynamics that generate conflict.
 James Liu and Denis Hilton (2005): How the past
weighs on the present: Social representations of history
and their role in identity politics.
 This study explores the aspects of historical
representation to explain the different responses by
America’s closest allies to the terrorist attacks of 9/11.
Comparative Historical Research:
 Timothy Snyder (2017): On Tyranny. Twenty Lessons
from the Twentieth Century.
 Taking a review on the autocratic rules of the
Twentieth Century, and making an attempt to point
out the similarities, the author argues that we must
learn from the horrors of the past if we want to protect
our democracy.
 He collects twenty suggestions, 20-point “how to” guide
for resisting tyranny.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Deductive-Analytic Method
Deductive-Analytic Method:
It is a research process during which general principles are
explored rested on analysing primary or secondary bibliography
that directly or indirectly related to the subject of the research.
 The challenge of this research design:
 Proper theoretic preparation
 To have appropriate knowledge about the related schools and theories,
their exponents, and the main theoretic debates.
 Analytic focus with interpretation:
To create models with highly abstract conceptualisation rested on analysing
the meanings of some key concepts and combining that analysis with logically
structured arguments showing the implications of a particular stance using
those concepts.
 Normative analysis provides a critical assessment of the assumptions and
philosophical foundations of political actions.
 It is worth starting with illustrative examples or presenting a case study,
then giving rational arguments for the acceptance of the general
principle.
 Focus not only on the pro but the con: try to find expectations and give
an explanation why they are out of the scope, or are irrelevant in the
discussion.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Case Study (#1)
Case Study:
The great advantage of the case study is
that by focusing on a single case, the
case can be intensively examined.
However, this is not enough.
Good case studies possess two important characteristics:
1. They say something interesting and meaningful about the case is being
studied.  The findings of the study should be internally valid.
 Example: A case study of ethnic violence in India should help to shed light
on the sources of conflict in India, and contribute to the academic
literature that has been written on the subject.
2. They should also aim to say more general, and engage with wider academic
debates that might be applicable to other contexts or other cases.  This
involves setting the case in comparative context, and proposing theories or
explanations that are externally valid (at least hipotetically).
 Example: Does the study only shed light on the origins of ethnic violence
in India, or could be implications of the argument and analysis also be
relevant for helping to explain the sources of ethnic conflict in other parts
of the world, such as in Africa?

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Case Study (#2)
Case Study Method:
Case Studies can be chosen to be
theory-confirming or theory-infirming
tools in research methodology.
 There are two main criteria for case selections:
1. The case should provide a fair test of the theory. That is to say, the case
should be representative of the domains of the theories they are intended
to test (the so-called paradigmatic case).
2. The case can be also used to examine specific outliers or deviant cases,
and to examine the cases (countries) that do not fit existing theory and are
known to deviate from established generalisations.
 All in all, we can distinguish between case studies that:
1. provide descriptive contextualisation;
2. apply existing theory to new contexts;
3. examine expectations to the rule;
4. generate new theory (in this case the selected paradigmatic case is called as
crucial case).

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Example: Cyberwarfare (#1)
Case Study: Russian-Estonian Conflict, 2007
Historic Background: The Bronze Soldier is the informal name of a
controversial Soviet World War II war memorial in Tallinn, the capital of
Estonia, which were relocated to the nearby Tallinn Military Cemetery.
The Russians consider the action as an offence of defamation.

Cyberattacks on Estonia:
 128 DDoS attacks;
 The length of time of each attack was between 2 and
10 hours;
 Techniques: DDoS:
 To spin up data traffic up to 100 MbpS (needed Distributed Denial of
huge zombie networks), which has reached the Service
1000-fold rate of the average;
 The Estonian Parliament had no internet
connection during 4 days;
 No Banking services in the country for more than
24 hours;
 Communication devices were inactive.
”The Parties agree that an armed attack against one or more of
them in Europe or North America shall be considered an attack
against them all” (from Article 5 of the Washington Treaty)
Nato Reply:
”4.a. Collective defence: NATO members will always assist each other against
attack, in accordance with Article 5 of the Washington Treaty. That commitment
remains firm and binding.”
Strategic Concept For the Defence and Security of NATO, Lisboa, 2010
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Example: Cyberwarfare (#2)
Generalisation: Is cyberwarfare a new generation of warfare?

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Comparative Methods
Comparative method is the most widely used
research design in social science.
 There are strict limits to real-life experiments in
socials; e.g., no labs for social researchers to
thorough observations about a regime transition or
a revolution.
 What we can really do is to compare facts, events
or processes that have been emerged in society.
 For effective comparisons, it is essential to create and apply an appropriate
conceptualisation, generally we need to use some typology or ideal type.
 The ultimate goal of typology or ideal type is to discover and express the
different forms of a certain phenomenon or a set of phenomena.
 Examples:
Political Regimes:
- Democracy
- Authoritarian Rules
- Hybrid Regimes

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Empirical-Qualitative Comparsion
Empirical-Qualitative Method
 It rests on small-N studies (and tending to select paradigmatic cases).
 The results of social researches cannot be expressed in ”hard data” but in
”soft data” involving case studies rather than numerical data.
Two Main Types of Empirical-Qualitative Method:
1. Most Similar System Design (MSSD)
Case 1 Case 2 Case N Selection is based on cases
that share many important
A A A Overall
characteristics but differ in
B B B one crucial respect related to
similarities the hypothesis of interest:
C C C
X  Y
X Not X Not X Crucial Independent Dependent
Y Not Y Not Y differences Variables

2. Most Different System Design (MDSD)


Selection is based on cases
Case 1 Case 2 Case N
that are different in most
respects and only similar on
A Not A Not A Overall the key independent variable,
B Not B Not B (X). The ultimate goal is if the
C Not C Not C differences hypothesis is supported, then
we should observe that our
X X X Crucial dependent variable (Y) is also
Y Y Y similarity similar across our cases.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Example: MSSD (#1)
 Research Question: What could be the key factors that might
account for the high-level of gun-related homicide in the US?
 Let’s use MSSD and make comparison Canada and the US.

Categories Canada US Comment


Per capita GDP $39,300 $47,000 Economic indicators
Unemployment rate 6.1% 7.2% (mainly similar)
Population below 10.8% 12%
poverty line
Income equlity (Gini 32.1 45 Significant variance?
index)
Urbanization 80% 82%
Overall
Religion Christian: 66% Christian:74% The majority is Catholic
in Canada and similarities
Protestant in the US
Ethnic Groups White: 66% White: 64%
Ethnic: 34% Ethnic: 36%
”Social indicators”
Literacy Rate 99% 99%
are similar

School attendance 17 years 16 years

Homicade rate: Higher in the US


Gun-related 0.54 2.97 by Crucial
Overall 1.75 5.75 550% differences
328%

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Example: MSSD (#2)
 Research Question: What could be the key factors that might
account for the high-level of gun-related homicide in the US?

What could be the crucial differences, i.e., the possible


independent variables, X:
- High level of gun ownership  pro: Finland; con: some
other countries with lower level of gun ownership, the
phenomenon exists.
- ”The culture of fear”
- Other: why can we perceive ”sinuous movements” in
the trend?

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Example: MDSD
 Theda Skocpol (1979): States and Social Revolutions: A
Comparative Analysis of France, Russia, and China. CUP.
 Argument in the book: In three different countries, Russia
(1917-21), France (1787-1800), and China (1911-49), social
revolutions occur when external military threats provoke a
split in the ruling elite and peasant communities –taking
advantage of this split– revolt.
 Hypothesis: Splitted elite  Revolution
 The author practically applies an MDS design.
France Russia China

A Not A Not A Overall


B Not B Not B
C Not C Not C differences
X – Elite Split X X Crucial
Y - Revolution Y Y similarity

A: Political System B: Social & Cultural Context C: Economic Conditions

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Qualitative methods

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Qualitative Methods: Summary
Qualitative Methods have both a descriptive and an analytical purpose:
 Descriptive Purpose is to try and provide accurate information about what
investigated actors think and do  Researcher tries to understand what
happens.
 Analytical purpose is to obtain ideas, data, and evidences for creating and
testing hypotheses and theories and understand why things happen.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Interviews
Types of (scientific) interviews:
 Standardised structured interviews
 Follow a common set of questions for each
interview.
 Ask the questions in exactly the same way, using
the same words, probes etc. for each interview.
 Present the participant with a set of answers to
choose from.
 Semi-structured interviews
 Follow a common set of topics or questions for each interview.
 May introduce the topics or questions in different ways or orders as
appropriate for each interview.
 Allow the participant to answer the questions or discuss the topic
in their own way using their own words.
 Unstructured interviews
 Focus on a broad area for discussion.
 Enable the participant to talk about the research topic in their own
way.
Semi-structured and unstructured interviews are regarded as ”non-
standardised.”
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Interview Guide: Example
Interview guide is an agenda Excuse me: we are carrying out a survey for the council to find out what people think
for an interview with additional about the park. Could you spare a few minutes to answer a few questions?

notes and features to aid the


researcher.

The following standardised


interview is being carried out
for the local council to find out
what users of the park think of
the park, and what changes
they would like to see.

Thank you for your help.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Interview Guide: Example (#2)
Semi-structured interview with women who were survivors of childhood sexual abuse
about their use of helping services.
1. Introduction 4. (cont.) Points to cover:
Chat about getting to interview etc. How are things going • Who did you tell?
with you today? • Can you describe what happened?
(Explain confidentiality – nobody else knows you are being • Can you remember why/how you came to tell that
interviewed – tapes will be destroyed.) person?
Interview will be a conversation. (Explain tape recorder and • What were you hoping would happen?
that she can switch it off at any time if not comfortable.) • What sort of help did you get?
2. Can you tell me a bit about yourself: how old you are, • What happened next?
where you live, who lives there with you? Children, etc.? 5. Subsequent use of services
3. Use of services Points as for current service.
Begin by asking about the service known to be used at 6. Good/not-so-good points of services used
present. 7. Check other service use
Points to cover: 8. Other support
• What sort of service? Help from family, partner, friends, other survivors?
• How often? 9. Future
• What happens there? Help needed in future? Type? How long?
• When did you first go? 10. Ending
• How did you hear about it? Any ways in which you would like services improved both
• Do the people there know you are a survivor? for yourself and for other survivors? Anything else that you
• How do you feel about going there? would like to say about the help survivors need?
4. First disclosure Switch off tape. Make sure participant is comfortable,
I want to ask you now about the first time as an adult you reassure about confidentiality and interest, chat, tea, etc.
told someone that you had been abused as a child.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Content Analysis
Content/Textual (Discourse) Analysis gives
opportunities to understand the backstory of events. It
makes the research be able to discover the deeper
context and/or hidden massages.
Textual Analysis can be applied as a stand-alone
research method or together with other methods.
 It can be the basis for both quantitative and
qualitative research frameworks.
 Content analysis involves the systematic inspection of
textual information.
 In politics and IR, researchers tend to study election
manifestos, news and social media, and political
leaders’ speeches.
 However, there is a wide variety of texts that
researchers might choose to analyse, including:
 Official documents: governmental and legal
reports, judicial decisions, company accounts;
records from law courts, schools, hospitals, etc.
 Cultural documents: newspaper articles and
editorials, magazines, TV programs, homepages,
etc.
 Personal documents: letters, diaries, blogs, and e-
mails.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Qualitative Data Analysis: Content Analysis

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Content Analysis: Coding Interviews

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Content Analysis: Coding Complex Contents

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Steps to carrying out content analysis
Step 1: To select the ”representative material” to be analysed.
 What set of documents is related to your research question?
 What sample from this set will you investigate?
Step 2: Define categories, key concepts or topics of interest that you will search
for in the material.
 You have two protocols to carry out Step1&2:
 Deductive: you create a vocabulary of categories, key concepts in
advance, and you analyse the text and classify the corpus of the text
accordingly. The task of the research program is to choose and rate the
appropriate classes in the vocabulary.
 Inductive: you are creating and extending the vocabulary of categories
and key concepts in the process of content analysis.
Step 3: To decide what segments of the text will contain what you are searching
for?
 Choose the recording unit of content:
 A single word or symbol
 A sentence or paragraph
 A theme
 A character
Step 4: Fix a coding protocol
 Coding the data
 Cleaning it up
 Classification: Develop a coding scheme and a thematic framework
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Coding Protocol (#1)
Coding the data: Lebel relevant words, phrases, sentences, or sections.
 Labels can be about persons, actions, activities, concepts, opinions,
differences, processes, or whatever you think is relevant.
 Something is relevant if
 it is repeated in several places of the text(s);
 the interviewee explicitly states/emphasis that it is important;
 it surprise you;
 You have read about something similar previously
 it reminds you of a theory or a concept.
Cleaning it up: Decide which codes are the most important and create
categories by bringing several codes together.
 Go through all the codes you have created.
 Read them and try to filter them, i.e., make attempts to create new
codes by combining two or more codes (reduce the number of codes).
 Keep the codes that you think are important and group them together
 Create and name categories
Classification: Decide which codes are the most relevant and how they are
connected to each other.
 Classify by categories
 (Optional) Classify by sources  for your Literature Review.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Coding Protocol (#2): Example
Adaptation

Seeking Info Problem Solving

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Coding Protocol (#3): Another Example

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Textual Analysis by Computer
Computer-based analysis:
 Qualitative content analysis: Word
cloud generator programs:
 [Link]
 [Link]
 [Link]

Example: US President Obama’s speech


in Cairo on June 4, 2009
”I've come here to Cairo to seek a new
beginning between the United States and
Muslims around the world,…”

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Textual Analysis by Computer
Computer-based analysis:
 Quantitative content analysis:
• [Link]: [Link]
• NVIVO: [Link]

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Validity and Reliability of CA
As a method of data collection, researchers using Content Analysis must be
concerned with the validity and reliability of their results.
 Validity: A study in CA is valid if
 its categories and explanations actually measures and explain what you
claim they measure and explain;
 your inferences follow from data;
 It is also important to see if
 We can draw unambiguous conclusions from our results, and
 Whether our conclusions are likely to apply to other similar situations
or cases.
 Reliability refers to ”consistency” and ”repeatability”:
1. Coder stability: does the same coder consistently recode the same data in
the same way over a period of time?
2. Reproducibility: do the coding schemes lead to the same text being coded
in the same category by two or more coders?
3. Intercoder reliability reveals ”objectivity” by showing the extent to which
different coders, each coding the same content, come to the same coding
decisions.
4. The main reasons for unreliability:
 The text itself is poorly written or vague.
 Word meanings, category definitions are ambiguous
 The coder makes mistakes

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Quantitative methods

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Dual Goal of Quantitaive Methods
The ‘statements’ which might be made on the basis of sample survey findings can be
descriptive, comparative, or relational:
 descriptive: ”10 per cent of adults play tennis”
 comparative: ”10 per cent play tennis, but 12 per cent play golf”
 relational: ”People with high incomes play tennis more than people
with low incomes.”
The dual goal of quantitative methods:
 Description: Summarizing the data for a particular variable.
 Inference: Using this information to make generalisations about the wider
population from which the sample was drawn.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Descriptive Statistics: Frequency Distribution (#1)
In descriptive statistics we can use tables, figures, and statistics.
Frequency Distribution – which describes the entire distribution of responses, and
summarises the number of cases for each given response code.
Example: We have a distribution of a sample of 150 people clustered by age and
gender. Start by finding out the distribution of the sample across the age groups by
gender.

This can also be presented in


a bar chart which illustrates
the different age distributions
of the men and women (and
total) in the sample.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Descriptive Statistics: Frequency Distribution (#2)
In descriptive statistics we can use tables, statistics, and figures.
Frequency Distribution – which describes the entire distribution of responses, and
summarises the number of cases for each given response code.
Example: We have a distribution of a sample of 150 people clustered by age and
gender. Start by finding out the distribution of the sample across the age groups by
gender.

Now, what if we wanted to know the


distribution of men and women in
each age group? The same table can
be used but the percentages are
worked out on the age groups (each
row) rather than each column
(gender).

Interesting finding: The youngest age


group is made up of 58.8% women
and the smallest group of men (14).
Why?

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Descriptive Statistics: Central Tendancy

There are other ways of summarising the data which may be used to look for
differences between different groups or categories.
A statistical measure which summarises the data relating to one variable in one
value, such as the mean, median or mode, is called measures of central tendency.
 The arithmetic mean is statistical average calculated by totalling all the values
and dividing by the number of cases.
 The median is a statistical average calculated by arranging all the values in a
sample in numerical order, then noting the middle value of the distribution.
 The mode is statistical average calculated by noting the most common value in
the distribution.
Example: There are eleven persons in a class, and each is asked to estimate their
total annual income (to the nearest thousand):

Note! While each can be useful in the early stages of analysis, it is important to also
look at the way in which the data is dispersed around the mean, median or mode –
so ensure that you take account of
 the shape of the distribution and
 the range of values (the gap between the lowest and highest values).
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Descriptive Statistics: Graphical Representation of Data (#1)

Histogram:
 The histogram is used to summarize data that are measured on an interval
scale, and to illustrate the major features of the distribution of the data.
 To use Excel’s Histogram data analysis tool, you must first establish a bin
array and then select the Histogram data analysis tool. In the dialog box
that is displayed, you then specify the input data (Input Range) and bin
array (Bin Range).

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Descriptive Statistics: Graphical Representation of Data (#2)

Box Plot:
 The box plot provides a pictorial
representation of the following statistics:
maximum, 75th-percentile, median (50th-
percentile), 25th-percentile, and minimum.
 As we shall see later, box plots are especially
useful when comparing samples and testing
whether or not data is symmetric.
 In Excel, we can create box plots directly:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Descriptive Statistics: Normal Distribution
Normal Distribution:
 In statistical terms a normal distribution is data that is distributed
symmetrically around the mean point in a ‘bell shape’. This is a bell-shaped
curve graph. Using statistical theory, we can say that if we took lots of
different samples from the same population, we would find normal distribution
(the so-called central limit theorem).

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Descriptive Statistics: Testing for Normality
Normality Tests:
 Normality tests are used to determine if a data set is well-modelled by a normal
distribution and to see if variables underlying the data set are normally
distributed.
 We have two approaches to assessing normality: Histogram Q-Q Plot
1. Visual Approach by using
o Histogram: The better fitting the data bars to
the bell curve, the closer is our data to normal
distribution.
o Q-Q Plot (Quantile-Quantile Plot) graphically
compares a sample of data on the vertical axis to
a statistical population on the horizontal axis.
The better fitting the plots constructed in this
way to the straight line, the closer is our data to
normal distribution.
2. Statistical Approach (not existed in Excel)
o If the result of K-S or S-W tests is significant
(Sig. < 0.05), then our data is NOT normal.
o The K-S and S-W tests are well-used on smaller
samples (n<30, <100). In the case of a larger
sample, we prefer to evaluate a Q-Q plot (or
check normality with a chi-square test.)
The result of K-S and S-W tests in SPSS

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Descriptive Statistics: Standard Deviation
Deviation from Normal Distribution:
 If we assume that the population mean is at the high point of the bell curve
(that is, coinciding with the highest numbers of sample means), then we can
say that our particular sample mean will lie within two standard deviations of
the central point in 90/95/99 per cent of cases.
 The standard deviation is a statistical measure of how the cases – in this
example the averages (means) of each of our samples – are distributed around
the mean (in this case the assumed population mean).

µ The standard deviation (SD) is


calculated by

 Note: if this is a sample and not the whole population, divide N-1.
(No the ’number of cases’ (N) but the ’degree of freedom’ (N-1) 
All the members of any series of numbers can be arbitrary,
except the last member that actually adds up to the sum. So, all
the data can be chosen free except the last one.)
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Descriptive Statistics: Standard Error and Confidential Interval

Standard Error:
 Suppose we know the mean height of the population is µ=169.
 We have six samples with the means s1=170.9, s2=168.3,…,s6=169.
 The standard error is the standard deviation of the means of the samples
s1,s2,…,s6.

Confidence Interval measures the degree of uncertainty or certainty in a sampling


method. They can take any number of probability limits, with the most common
being a 95% or 99% confidence level.
 A confidence interval is a range of values, bounded
above and below the mean.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The 95% Rule
MEAN
95% RULE
If a distribution of data is
ONE STANDARD approximately bell-shaped,
DEVIATION
about 95% of data fall within
SD = two standard deviations of
the mean.

TWO STANDARD
DEVIATIONS

Z = -1,96 Z = 1,96 Reliability: How precise is


our stat?
At 95% we have only 5
95% CONFIDENCE
INTERVAL
errors out of 100 score.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Samlple Size
How to determine the sample size?
 Calculate the sample size for any (”infinite”) population – Cochran formula
 Adjust the sample size to required population
Cochran Formula
 Sample size (S) = x p x (1 – p)/ where
 z: z score = 1.96 at 95% confidence level (1.65 at 90%, and 2.33 at 99% confidence
level)
 p: population proportion (p = 0.5 is assumed in the general case)
 ME: margin of error that is a small amount that is allowed for in case of miscalculation
or change of circumstances (generally we also set ME = 0.05).
 In general case (”infinite population”) the Sample size = x 0.5 x 0.5/  385.
Example: A sporting goods company wants to estimate the proportion of tennis players among
college students with a fairly high degree of accuracy. They want to get an estimate with a
precision of 0.02 at a 99% confidence interval. A pilot survey by telephone among 100 high school
students showed that 23 out of them played tennis. Estimate the sample size required for the
complete research.
 z score = 2.33 at 99% confidence level, p=23/100, ME=0,02
 The Sample size = x 0.23 x 0.77/  2404 (2403,645).
Adjusted Sample Size:
 Adjusted Sample size = S/1+ [(S – 1)/N] where
 S is the general sample size calculated by Cochran formula
 N is the size of the required population
 Suppose the population is 570.000 (the number of college students).
Then, the adjusted sample size is 2404/1+[(2404 – 1)/570.000]  2404 (2404,004 ).
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Statistical Hypotheses
A common feature of the statistical method is the concept of the null hypothesis,
referred to by the symbol H0.
 It is based on the idea of setting up two mutually incompatible hypotheses, so
that only one can be true.
 Example: Either more people play tennis than golf or the number of people
who play tennis is less than or equal to the number who play golf – if one
proposition is true then the other is untrue.
The null hypothesis usually proposes that there is no difference between two observed
values or that there is no relationship between variables. There are therefore two
possibilities:
H0 – Null hypothesis: there is no
significant difference or relationship
H1 – Alternative hypothesis: there
is a significant difference or
relationship.
Note! Usually it is the alternative hypothesis,
H1, that the researcher is interested in, but
statistical theory explores the implications of
the Null hypothesis.

In terms of Research Methodology, this is very much a deductive approach: the


hypothesis is set up in advance of the analysis, possibly within a theoretical
framework.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Statistical Tests
The different statistical tests are associated with different levels of measurement.
 The tests all relate to comparisons between variables and relationships between
variables.
For testing
Associational
hypotheses Nominal and
(1st Type of Ordinal Statistics
Hypotheses)

Interval Statistics

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Statistical Tests
The different statistical tests are associated with different levels of measurement.
 The tests all relate to comparisons between variables and relationships between
variables.
For testing
variance-type
hypotheses
(2nd Type of
Hypotheses)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Patterns of Association

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Statistical Tests
The different statistical tests are associated with different levels of measurement.
 The tests all relate to comparisons between variables and relationships between
variables.
For testing
Associational
hypotheses Nominal and
(1st Type of Ordinal Statistics
Hypotheses)

Interval Statistics

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Cross-Tabulation and Chi-Square Test (#1)
The Chi-squared is a test that can be used to assess whether the difference
between the mean values of two samples is statistically significant (by significant
we mean worthy of consideration and note). It can be used to compare variables
that are nominal or ordinal, for example, male and female.
We will use it to test the statistical significance of the data in the cross-tabulation.
Example: 100 adults, men and women, are asked how much money (EUR) they
spend per month on cosmetics. Here is a chart summing up the facts:

 To test for statistical significance, we set up a null hypothesis that there is no


relationship between gender and the spending on cosmetics – that the two variables are
independent of each other.
 The chi-squared test is based on measuring how far the observed values (the data
collected) differ from those that would be expected if the two groups (in this case men and
women) were the same in terms of the variable ’spending on cosmetics.’
 H0 is that there will be no significance difference between males’ and females’
spending on cosmetics.
 H1 is that in this case the difference between men and women is statistically
significant

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Cross-Tabulation and Chi-Square Test (#2)
Example: 100 adults, men and women, are asked how much money (EUR) they spend per
month on cosmetics. Here can be seen the fact (observed) and the expected values:

 The calculation is for where observed values are your data.


STEPS:
1. Each expected value is worked out by (column total X row total)/grand total, e.g.,
column total for under 25 (42) x row total for male (40)/grand total (100) = 16,8.
2. We can calculate the Chi-Square Value in our example:

3. The result of the calculation is adjusted according to size of the table – the number of
categories in each of the variables included (spending has 3 categories and gender has
2). From this, the degrees of freedom are worked out. In this case there are
(3 - 1)(2 - 1) = 2 degrees of freedom.
4. By using an Excel function, we can calculate the critical Chi-Square value:
[Link](alpha; the degree of freedom) = [Link](0,05;2), as the confidence
level is 95%, so alpha =0,05. The critical value (CHISQcritical) is 5,9915.
 c2 is than the critical value, which means that we have good reason to reject H 0. Otherwise,
if CHISQcritical > c2 , we may accept the Null Hypothesis, H0.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Regression: What is it?
We now look at statistics that evaluate the relationship between interval variables
regression. These statistics are derived from a procedure called regression.
Regression statistical models describe the relationship between variables by
fitting a line to the observed data.
 Linear regression models use a straight line, while nonlinear regression
models use a curved line.
 Regression allows you to estimate how a dependent variable changes as
the independent variable(s) change.
Simple linear regression is used to estimate the relationship between the two
types of variables.
 To display the data in a Cartesian system, called Scattergram, we have some
points, and want to have a line that best fits them like this:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Linear Regression
Linear Regression is a statistical models describe the relationship between
variables by fitting a line to the observed data.
 We can place the line "by eye": try to have the line, and call trendline, as
close as possible to all points.
 Some points will be above and some others will be below the line, and the
distances between the points and the trendline are the “errors,” and the
points are called outliners.

Prerequisites for applying Linear Regression:


1. Two variables should be measured at the continuous level (i.e., they are
either interval or ratio variables).
2. There needs to be a linear relationship between the two variables.
3. There should be no significant outliers.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
The Conception of Linear Regression
The math of Linear regression:

The Linear Regression model postulates that two random variables X and Y are
related by a straight line as follows:

Y = a + bX

where

Y is the dependent variable


X is the independent variable
a is the intercept
b is the slope

The math of Linear regression:


Linear Regression estimates the coefficients of the linear equation, involving one
independent variables that best predict the value of the dependent variable.
 Example:
The number of cigarettes per day (independent) and life expectancy
(dependent): Y =a+b  X
Life expectancy = a + b  Number of cigarettes
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
The Least Squares Method
We want to minimalise the sum of the distances between the points and the line. We have two
ways to minimise:
1. First regression line: We consider the “vertical distances” (parallel with Y axis) between
the points and the trendline, and then we try to set the minimise of the sum of the distances.

1. ”Vertical distance” 2. ”Horizontal distance”

2. Second regression line: We consider the “horizontal distances” (parallel with X axis)
between the points and the trendline, and then we try to set the minimise of the sum of the
distances.
However, these distances can be positive or negative due to the position of the points, whether
they are above or below the trendline. We can exclude this problem if we take the square of the
values of the distances.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Linear Regression in action (#1)
Example: The table below shows the annual income (yi) and the number of
employees (xi) of the World top 12 companies in vehicle industry.
 Question: Is there any relationship between these two variables, i.e. the
companies’ annual income and how many employees do they hire?
We analyse the data by calculating the first and second regression lines:

xi yi xi y i xi^2 yi^2
756,3 123,8 93629,94 571989,69 15326,44
332,7 89 29610,3 110689,29 7921
102,4 78,1 7997,44 10485,76 6099,61
379,3 57,3 21733,89 143868,49 3283,29
287,9 46,8 13473,72 82886,41 2190,24
265,6 46 12217,6 70543,36 2116
138,2 42,9 5928,78 19099,24 1840,41
85,5 30,6 2616,3 7310,25 936,36
147,2 29,4 4327,68 21667,84 864,36
126,5 29,4 3719,1 16002,25 864,36 first: Income(y) = 0,13 employees(x) + 20,882
159,1 29,3 4661,63 25312,81 858,49
156,8 28,4 4453,12 24586,24 806,56 second: Income(y) = 0,24 employees(x) - 5,29
2937,5 631 204373,79 1 104 469,28 43107,12
xi: the number of employees (in thousand) This calculation is done "behind the scenes" by the
computer. The calculation is not important, follow the logic!
yi: income (in billion $)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Linear Regression in action (#2)
Here you can see the first regression line in red, and the second in green.
 Usually the first regression lines tend to bend to X axis as we minimalise the
distances subject to y.
 Similarly, the second lines tend to bend to Y axis as we minimalise the
distances subject to x.
The required trendline (in blue colour)
can be determined by using computer
softwares (Excel, SPSS, R).
 In Excel, e.g., the method is as
follows:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Correlation
Parson’s R: In statistics, we are talking about the strength and direction of the
relationship of correlation in accordance with the value of the correlation coefficient R.

Negative correlation

In the above example, we have


Income(y) = 0,13  employees(x) + 20,882

and R=0,81 that shows a strong correlation.

Another important factor is the coefficient of


determination (R²) which is the square of Parson’s
R, and it measures how well the statistical model Positive correalation
predicts the outcome. SUMMARY OUTPUT
 As the outcome is represented by the
dependant variable, and the explanatory part Regression Statistics
Multiple R 0,806874197 Pearson’s R
of the prediction is based on the independent
R Square 0,65104597 Coeff, of Determ.
variable, we often say that R² indicates that how Adjusted R Square 0,616150567
much the dependent variable (y) is explained (or Standard Error 18,61203763
predicted) by the independent (predictor) variable
The Excel Summary Output of the Regression
(x) in the linear regression model.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Interpret the result of the Linear Regression
Step 1: Run the Regression test Data Table xi yi xi y i xi^2 yi^2
Press Data > Analysis | Data Analysis and 756,3 123,8 93629,94 571989,69 15326,44
332,7 89 29610,3 110689,29 7921
selecting Regression from the menu that is
102,4 78,1 7997,44 10485,76 6099,61
displayed: 379,3 57,3 21733,89 143868,49 3283,29
287,9 46,8 13473,72 82886,41 2190,24
265,6 46 12217,6 70543,36 2116
138,2 42,9 5928,78 19099,24 1840,41
85,5 30,6 2616,3 7310,25 936,36
147,2 29,4 4327,68 21667,84 864,36
126,5 29,4 3719,1 16002,25 864,36
159,1 29,3 4661,63 25312,81 858,49
156,8 28,4 4453,12 24586,24 806,56
2937,5 631 204373,79 1 104 469,28 43107,12
xi: the number of employees (in thousand)
SUMMARY OUTPUT yi: income (in billion $)
Step 2: Interpret the result
The regression model is given by the Regression Statistics
The Excel outcome of
Multiple R 0,806874197
formula: income = xi  employees + b R Square 0,65104597 the Regression
Adjusted R Square 0,616150567
Income(y) = 0,13  employees(x) + 20,882 Standard Error 18,61203763
Observations 12
• The p-value of the regression is 0.005
(highlighted by green) which is below ANOVA
0.05, so we reject the null hypothesis. Regression
df
1
SS
6462,957219
MS
6462,957
F
18,65707
Significance F
0,001514586
• R square indicates that 65% of the Residual 10 3464,079448 346,4079
variability in the income can be Total 11 9927,036667

explained (predicted) by differences in Coefficients Standard Error t Stat P-value


the number of employees. Intercept (b) 20,8821473 9,095737619 2,295817 0,04457
xi 0,129502717 0,029981763 4,319383 0,001515

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Reliability of Linear Regression
The most important parts of the output of a
Linear Regression done in Excel are:
1. Overall Regression Accuracy (correlation
test): how much the independent
variable explains the occurrence of
phenomenon associated with the
dependent variable (R Square).
2. Probability that the regression is not due
to random: significance of F should be
below 0.05.
3. Reliability of the regression’s Y-intercept
and coefficients (and the p-values of Y-
intercept and coefficients) that gives the
formula of the model:
Y = a (0,1295)  X + b (20,882)
4. Residuals show now patterns: if we plot 50

the residual (actual – predicted values of 40


30
the model) and the points show a curved
Residuals

20
(e.g., U-shaped) pattern, a non-linear 10
model might probability fit better. 0
-10 0 100 200 300 400 500 600 700 800
-20
X Variable 1

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Effect Size
Effect size always indicates the practical significance of a research outcome. A
large effect size means that a research finding has practical significance, while a
small effect size indicates limited practical applications.
 Under associational (first type of) hypothesis, effect size tells you how strong
the relationship between variables is.

For Chi-Square test, we use the Cramer’s Phi or V’s value:

where
c2 is the calculated Chi-squared value, N is the size of the sample, r and c are the rows and
columns of the cross-table.
The strengh of the effect size is found in the table below:

For correlation, we can use either Person’s R or the coefficient of determination


(R²) with the same interpretation as we studied them beforehand.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


(Variance-type) Hypothesis Testing

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Statistical Tests
The different statistical tests are associated with different levels of measurement.
 The tests all relate to comparisons between variables and relationships between
variables.
For testing
variance-type
hypotheses
(2nd Type of
Hypotheses)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


One Sample t-test: Introduction
This is a test for difference between sample mean and pre-determined population
mean.
Examples:
 Comparison of mean dietary intake of a particular group of individuals with the
recommended daily intake.
 Average birth weight of new born baby in Malaysia
Hypotheses in one sample t-test:
 2nd (variance-type) hypothesis, i.e., some variables have been changed because
of an intervention (experiment); the question is if the change was due to the
intervention or chance.
 Statistically we have
o H0 – Null hypothesis: There is no significant difference between the sample
mean and the population mean.
o H1 – Alternative hypothesis: there is a significant difference between the
sample mean and the population mean.
Requirements for applying one sample t-test (the reliability of the test):
1. Dependent variable should be measured at the interval or ratio level.
2. There should be no significant outliers.
3. Dependent variable should be approximately normally distributed.
4. The data are independent, i.e., there is no relationship between the
observations.
Important Note! Requirements #1–#3 are shared by all the different t-tests.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
The Reliability of t-tests
Prerequisites for applying t-tests is a test for normality:
 As requirements #1–#3 are shared by all the different t-tests, we have got to use one of the
normality tests (e.g., Q-Q plot) before applying any t-test.
Or the sample data are reasonably symmetric related to the sample mean:
 The t-distribution provides good results even if
• the population is not normal, and
• the sample is small,
provided the sample data is reasonably symmetrically distributed about the sample mean.
 The following attributes are good indications of symmetry:
o The box plot is relatively symmetrical; i.e., the median is in the center of the box and
the whiskers extend equally in each direction;
o The histogram looks symmetrical;
o The mean is approximately equal to the median;
o The coefficient of skewness is relatively small.

A reasonably symmetric box plot A non-symmetric box plot

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


One Sample t-test
Example: A weight reduction program claims to be effective in treating obesity. To
test this claim, 12 people were put on the program and the number of kilogramms of
weight gained/lost was recorded for each person after two years.
Step 1: Test for Normality
Box plot of the sample data is quite
symmetric (see previous slide) about
the sample mean (5.5)
Step 2: Set up the Hypothesis
H0: The weight reduction program is
not effective;
H1: The program is effective.
Step 3: Doing the test
- As p value > alpha  H0 is plausible
to accept at 95% level
-As t < t critical  We reject H1 to
accept at 95% level
Step 4: Public the findings
We made a one-sample t-test with
Findings: values t=1,062; df=11; p=0,311, thus
it can be concluded that we cannot
t df Sig. Mean diff. reject the null hypothesis. All in all,
Weight Loss 1,062 11 0,311 5,500 the program is not effective at a
statistically significant level (CI=95%).

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Two-Samples Independent t-test (#1)
When and how to use?
 Our goal is to determine whether or not the means of two populations are equal given two
independent samples, one from each population.
 The test has two forms in accordance of equal variance () assumption:
o 1  2: Assuming that population variances are equal even if unknown. (Thumb of
rule: generally, even if one variance is up to four times the other, the equal variance
assumption will give good results.)
o 1 ≠ 2: Assuming unequal population variances (see Example in the Seminar Notes).
Example: A food company wants to determine whether or not their new formula for peanut
butter is significantly tastier than their old formula. They chose two random samples of 10 people
and ask the people in the first sample to taste the peanut butter using the existing formula, and
they ask the people in the second sample to taste the peanut butter with the new formula. They
then ask all 20 people to fill out a questionnaire rating the tastiness of the peanut butter they
tasted. Based on the data, determine whether there is a significant difference between the two
types of peanut butter.
Step 1: Test for Normality
Box plot of the sample data is
reasonably symmetric.
Step 2: Verifying the variances
The variances for the two samples
are 13.3 and 18.8. The two values
are close enough to satisfy the
equal variance assumption:
1  2
Note: A more exact test is Levene’s f-test.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Two-Samples Independent t-test (#2)
Example: A food company wants to determine whether or not their new formula for peanut
butter is significantly tastier than their old formula. They chose two random samples of 10 people
and ask the people in the first sample to taste the peanut butter using the existing formula, and
they ask the people in the second sample to taste the peanut butter with the new formula. They
then ask all 20 people to fill out a questionnaire rating the tastiness of the peanut butter they
tasted. Based on the data, determine whether there is a significant difference between the two
types of peanut butter.
Step 3: Set up the Hypothesis
H0: The difference between the mean scores for the two
formula is due to random effects.
t-Test: Two-Sample Assuming Equal Variances
H1: There is a statistically significant difference.
Step 4: Doing the test in Excel NEW OLD
Use Excel’s t-Test: Two Sample Assuming Equal Variances data analysis tool. Mean 15 11,1
 We begin by selecting Data > Analysis|Data Analysis. Variance 13,333333 18,76667
 When the dialog box appears, you need to fill in the values: Observations 10 10
Pooled Variance 16,05
Hypothesized Mean
Difference 0
df 18
t Stat 2,1767677
P(T<=t) one-tail 0,0215264
t Critical one-tail 1,7340636
P(T<=t) two-tail 0,0430527
t Critical two-tail 2,100922
 The output of the analysis is displayed on the right side.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Two-Samples Independent t-test (#3)
Example: A food company wants to determine whether or not their new formula for peanut
butter is significantly tastier than their old formula. They chose two random samples of 10 people
and ask the people in the first sample to taste the peanut butter using the existing formula, and
they ask the people in the second sample to taste the peanut butter with the new formula. They
then ask all 20 people to fill out a questionnaire rating the tastiness of the peanut butter they
tasted. Based on the data, determine whether there is a significant difference between the two
types of peanut butter.
t-Test: Two-Sample Assuming Equal Variances Step 5: Public the findings
• Group means are significantly
NEW OLD
different because p value
Mean 15 11,1
(p=0,043) is less than 0.05.
Variance 13,333333 18,76667 • The means are: 15 and 11,1.
Observations 10 10 • Std. Deviations (the square root
Pooled Variance 16,05
Hypothesized
of variances): 3,65 and 4,33.
Mean Difference 0
df 18 Report:
t Stat 2,1767677 This study found significantly
P(T<=t) one-tail 0,0215264
t Critical one-tail 1,7340636
larger values (15 ± 3.65) of the
P(T<=t) two-tail 0,0430527 new formula compared to the old
t Critical two-tail 2,100922 one (11 ± 4.33), t(18) = 2.101, p =
0.043.

Findings: Group N Mean Std. Dev.


New 10 15 3,651 t-test for Mean
OLD 10 11,1 4,332 equlity of t df Sig. diff.
means 2,101 18 0,043 3,9

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Paired Samples t-test (#1)
When and how to use?
 Paired Samples t-test is used if two related samples differ significantly from one another.
 Unlike the hypothesis testing studied so far, the two samples are not independent of one
another.
 This test always related to time-series researches, i.e., same individuals are studied more
than once in different time.
Example: A clinic provides a program to help their clients lose weight and asks a consumer
agency to investigate the effectiveness of the program. The agency takes a sample of 15 people,
weighs each person in the sample before the program begins, and then weighs each person three
months later to produce the results in the table below. Determine whether the program is
effective.

Step 1: Test for Normality


Box plot of the sample data is
quite symmetric

Step 2: Set up the


Hypothesis
H0: The weight reduction
program is not effective, the
differences in weight is due to
chance;
H1: The program is effective.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Paired Samples t-test (#2)
Example: A clinic provides a program to help their clients lose weight and asks a consumer
agency to investigate the effectiveness of the program. The agency takes a sample of 15 people,
weighs each person in the sample before the program begins, and then weighs each person three
months later to produce the results in the table below. Determine whether the program is
effective.

Step 3: Doing the test in Excel


t-Test: Paired Two Sample for
Use Excel’s t-Test: Paired Two Sample for Means data analysis tool. Means
 We begin by selecting Data > Analysis|Data Analysis.
 When the dialog box appears, you need to fill in the values: Before After
Mean 207,9333333 197
Variance 815,7809524 595
Observations 15 15
Pearson Correlation 0,983720406
Hypothesized Mean
Difference 0
df 14
t Stat 6,689699535
P(T<=t) one-tail 5,13783E-06
t Critical one-tail 1,761310136
P(T<=t) two-tail 1,02757E-05
t Critical two-tail 2,144786688

 The output of the analysis is displayed on the right side. The Excel output of the Paired sample
t-test analysis

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Paired Samples t-test (#3)
Example: A clinic provides a program to help their clients lose weight and asks a consumer
agency to investigate the effectiveness of the program. The agency takes a sample of 15 people,
weighs each person in the sample before the program begins, and then weighs each person three
months later to produce the results in the table below. Determine whether the program is
effective.
t-Test: Paired Two Sample for
Step 4: Public the
Means
findings
Before After Report: the values of t(df) and
Mean 207,9333333 197 p=sign. level
Variance 815,7809524 595
Interpretation:
Observations 15 15
We made the pair-dependant t-
Pearson Correlation 0,983720406
test with values t(14)= 6,690,
Hypothesized Mean
p < 0,05. As the test is proved
Difference 0
statistically significant, H1 is
df 14
accepted. The mean afterward is
t Stat 6,689699535 lower than the mean weight
P(T<=t) one-tail 5,13783E-06 beforehand, we can conclude
t Critical one-tail 1,761310136 there is a significant reduction in
P(T<=t) two-tail 1,02757E-05 weight. Thus, it is fair to say that
t Critical two-tail 2,144786688 the program is effective.

Findings: Mean
t df Sig. diff.
Pair 6,690 14 0,000 10,933

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


One-Way ANOVA: Introduction
When and how to use?
 Essentially, the analysis of variance (ANOVA) is an extension of independant
samples t-test.
 The one-way ANOVA is used to determine whether there are any significant
differences between the means of three or more independent (unrelated) groups.
Example:
 A one-way ANOVA is used to understand whether exam performance differed based on test
anxiety levels amongst students, dividing students into three independent groups (e.g., low,
medium and high-stressed students).
 A food company wants to determine whether or not any of
their three new formulas for peanut butter is significantly
tastier than their old formula. They create four random
samples of 10 people each (one for each type of peanut
butter) and ask the people in each sample to taste the
peanut butter for that sample. They then ask all 40 people
to fill out a questionnaire rating the tastiness of the peanut
butter they tried. Based on the data on the right,
determine whether there is a significant difference between
the four types of peanut butter.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


The Reliability of One-Way ANOVA
Why cannot we do multiple t-tests?
• First off, we can use a t-test to determine whether or not
there is a significant difference between the rating for
New 2 and New 1 (see the t-test result on the right). And
we get p=0.057>0.05, so there is no significant difference
between New 2 and New 1.
• If we instead compare the New 2 with the OLD formula,
we get a statistically significant result: p= 0.004 < 0.05.
• We can continue these pairwise comparisons, and draw a
conclusion, but it may imply an experimentwise error.
The problem with this approach is that doing multiple
tests, we will essentially increase our overall type I error
which can be higher than we would like.
 In fact, when we use a significance level of α = 0.05, we accept that five percent of the time we will get a
type I error. If we perform just three such tests, then we will essentially increase our overall type I error to
1 – (1 – 0.05) = 0.14. This means that 14 percent of the time, we will have a type I error, which is too high
to trust in the result.

Requirements for applying ANOVA:


1. Dependent variable should be measured at the interval or ratio level.
2. The independant vareiable should consist of two or more categorical, independant
groups.
3. There should be no significant outliers.
4. Dependent variable should be approximately normally distributed for each category
of the independant variable.
5. The data are independent, i.e., there is no relationship between the observations.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
One-Way ANOVA in action (#1)
Example: A food company wants to determine whether or not any of their three new formulas for
peanut butter is significantly tastier than their old formula. They create four random samples of
10 people each (one for each type of peanut butter) and ask the people in each sample to taste the
peanut butter for that sample. They then ask all 40 people to fill out a questionnaire rating the
tastiness of the peanut butter they tried. Based on the data below, determine whether there is a
significant difference between the four types of peanut butter:

Step 1: Test for Normality


Box plot of the sample data is
reasonably symmetric

Step 2: Set up the


Hypothesis
H0: Any difference between
the four types of peanut
butter is due to chance;
H1: There is a significant
difference between the four
types of peanut butter.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


One-Way ANOVA in action (#2)
Example: A food company wants to determine whether or not any of their three new formulas for
peanut butter is significantly tastier than their old formula. They create four random samples of
10 people each (one for each type of peanut butter) and ask the people in each sample to taste the
peanut butter for that sample. They then ask all 40 people to fill out a questionnaire rating the
tastiness of the peanut butter they tried. Based on the data below, determine whether there is a
significant difference between the four types of peanut butter:

Step 3: Doing the test in Excel


Anova: Single Factor
Use the ANOVA: Single Factor data analysis tool.
 To access this tool, press Data > Analysis|Data SUMMARY
Groups Count Sum Average Variance Std. Dev.
Analysis and fill in the dialog box that appears:
OLD 10 111 11,1 18,76667 4,33
NEW1 10 131 13,1 21,21111
NEW2 10 166 16,6 7,155556 2,68
NEW3 10 119 11,9 26,32222

ANOVA
Source of
Variation SS df MS F P-value F crit
Between Groups 176,675 3 58,89167 3,206928 0,034463 2,866266
Within Groups 661,1 36 18,36389

Total 837,775 39

The Excel output of the ANOVA: Single-


 The output of the analysis is displayed on factor data analysis
the right side.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
One-Way ANOVA in action (#3)
Example: A food company wants to determine whether or not any of their three new formulas for
peanut butter is significantly tastier than their old formula. They create four random samples of
10 people each (one for each type of peanut butter) and ask the people in each sample to taste the
peanut butter for that sample. They then ask all 40 people to fill out a questionnaire rating the
tastiness of the peanut butter they tried. Based on the data below, determine whether there is a
significant difference between the four types of peanut butter:

Step 4: Public the findings ANOVA


Source of Variation SS df MS F P-value F crit
 We can see in the ANOVA table that
Between Groups 176,675 3 58,89167 3,206928 0,034463 2,866266
the significance level is 0.034, which
Within Groups 661,1 36 18,36389
is below 0.05. So we can accept H1,
i.e., here is a significant difference Total 837,775 39
between the four types of peanut
butter. Report: There was a statistically significant difference
Step 5: Post hoc (Follow-up) Analysis between groups as determined by one-way ANOVA
 Although we now know that there is (F(3,36) = 3.207, p = 0.034).
a significant difference between the (Using the Post hoc results (see the next slide)):
four types of peanut butter, we still Tukey HSD post hoc test revealed that the New2
don’t know where the differences lay. formula was statistically significantly tastier (16.6
 Post hoc range tests and pairwise ± 2.68, p = 0.033) compared to the old one (11.1 ±
multiple comparisons can determine 4.33). The Bonferonni test verifies this (p=0.041).
which means are differ but this is However, it turns out there were no statistically
not supported by the Excel. (See the significant differences related to the other formulas.
next slide, too)

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


On Post hoc Tests
The Idea of post hoc tests:
 Although we now know that there is a
significant difference between the four
types of peanut butter, we still don’t
know where the differences lay.
 We have clarified that we cannot do
multiple t-tests for this purpose,
because it may imply an experimentwise
error.
 The general approach for addressing
this issue is to use different types of
follow-up tests (e.g. Bonferoni, Tukey’s
HSD, etc.) for carefully find the crucial
factor with keeping down the degree of
making Type I error.
 In SPSS, we can do these tests, and the
rusult is displayed on the right:
 We have significant mean
difference between OLD and New2
formula where calculated p is
below 0.05.

The SPSS output of the ANOVA: Post hoc tests

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Effect Size
Effect size always indicates the practical significance of a research outcome. A
large effect size means that a research finding has practical significance, while a
small effect size indicates limited practical applications.
 Under variance-type (second type of) hypothesis, effect size tells you how
strong the difference between groups is.
For t-tests and ANOVA, we use the following formulae:
 Cohen’s d:

o Paired t-test:

o Independent t-test:

 ANOVA: and
effect: between groups
error: within groups
The strengh of the effect size is found in the table below: total: total

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


A Brief Guide to Multivariate Analysis

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Statistical Tests
The different statistical tests are associated with different levels of measurement.
 The tests all relate to comparisons between variables and relationships between
variables.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Multiple Regression: Introduction
Multiple regression is linear regression involving more than one
independent variable.
 Example:
1) We might hypothesis that sports participation is dependent not
just on income but also on age,
Sports participation (y) = a + b × income (x1) + c × age (x2).
2) Overseas travel is dependent not just on income but also on the
price of airfares:
Travel (y) = a + b × income (x1) + c × fares (x2).

It is possible, in theory, to continue to add variables to the equation. This


should, however, be done with caution, since it frequently involves multi-
collinearity, where the independent variables are themselves inter-
correlated.
 Often, in leisure and tourism or in politics, a large number of variables is
involved, many inter-correlated, but each contributing something to the
phenomenon under investigation.
 The ‘independent’ variables should be, as far as possible, just that:
independent. Various tests exist to test for this phenomenon.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Multiple Regression in action (#1)
Example: A jeweler prices diamonds on the basis of quality (with values from 0 to 8, with 8 being
flawless and 0 containing numerous imperfections) and color (with values from 1 to 10, with 10
being pure white and 1 being yellow). Based on the price per carat of the 11 diamonds, each
weighing between 1.0 and 1.5 carats as shown in the table below, build a regression model which
captures the relationship between quality, color, and price.

Step 1: Run the Regression test


Press Data > Analysis | Data Analysis and
selecting Regression from the menu that is
displayed:

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Multiple Regression in action (#2)
Example: A jeweler prices diamonds on the basis of quality (with values from 0 to 8, with 8 being
flawless and 0 containing numerous imperfections) and color (with values from 1 to 10, with 10
being pure white and 1 being yellow). Based on the price per carat of the 11 diamonds, each
weighing between 1.0 and 1.5 carats as shown in the table below, build a regression model which
captures the relationship between quality, color, and price.

Step 2: Interpret the result


Adjusted R Square is an attempt to One of the key results of the
create a better estimate of the true analysis is that the regression
value of the population coefficient of model is given by the formula:
determination, and is calculated by
the formula: 1-(1-R2)x(N-1)/N-df-1, Price = 1.7514 + 4.8953 x Color +
i.e., 1-(1-B5)*(B8-1)/(B8-B12-1) + 3.7584 x Quality.

• The p-value of the regression is 0.005


(cell F12) which is below 0.05, so we
reject the null hypothesis.
• Color and Quality coefficients (cell
E18 and E19) are significant (<.05),
whilst the intercept coeff. is not.
• Adjusted R square shows strong
correlation between the price and both
the color and quality.

 Based on the data, we can estimate the price if the diamond has color 4 and quality 6:
Price = 1.75 + 4.90 ∙ Color + 3.76 ∙ Quality = 1.75 + 4.90 ∙ 4 + 3.76 ∙ 6 = 43.88

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Factorial ANOVA: Introduction
When and how to use?
 We now extend the one-factor ANOVA model previously described to a model with more
than one factor.
 A factor is an independent variable  Practically, Factorial ANOVA is a multivariate
analysis.
We have
 A level is some aspect of a factor.  two factors: - Blends, and – Crops;
 Let’s look at a two-factor model now.  three levels of Blend: X, Y, and Z;
 The two-way ANOVA compares the mean  three levels of Crops: Wheat, Corn, Soy,
and Rice.
differences between groups that have been
split on two independent variables (called
factors).
 The primary purpose of a two-way ANOVA is
to understand if there is an interaction between
the two independent variables on the dependent
variable.
Example:
A new fertilizer has been developed to increase the
yield on crops. The makers of the fertilizer want to
better understand which of the three formulations
(or blends) of this fertilizer are most effective for
wheat, corn, soy beans, and rice (the crops). They
test each of the three blends on five samples of
each of the four types of crops. The crop yields for
the five samples of each of the 12 combinations are
as shown in the table on the right.
Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY
Two-Way ANOVA in action (#1)
Example: A new fertilizer has been developed to increase the yield on crops. The makers of the
fertilizer want to better understand which of the three formulations (or blends) of this fertilizer are
most effective for wheat, corn, soy beans, and rice (the crops). They test each of the three blends
on five samples of each of the four types of crops. The crop yields for the five samples of each of
the 12 combinations are as shown in the table below.
Step 1: Set up Hypotheses to test
 H0 for Factor A: There are no significant differences between the
blends.
 H0 for Factor B: There are no significant differences between the
effectiveness of the fertilizer for the different crops.
 H0 for Interactions: There are no significant differences between
crop and blend.
Step 2: Run the two-way ANOVA test
Use the ANOVA: Two Factor with Replication data analysis
tool. To access this tool, as usual, Data > Analysis|Data
Analysis and fill in the dialog box:

We will get
this output

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Two-Way ANOVA in action (#2)
Example: A new fertilizer has been developed to increase the yield on crops. The makers of the
fertilizer want to better understand which of the three formulations (or blends) of this fertilizer are
most effective for wheat, corn, soy beans, and rice (the crops). They test each of the three blends
on five samples of each of the four types of crops. The crop yields for the five samples of each of
the 12 combinations are as shown in the table below.

Anova: Two-Factor With Replication Step 3: Public the findings


 Since the p-value (blends) = 0.00025 < 0.05 = α, we reject the
SUMMARY Wheat Corn Soy Rice Total Factor A null hypothesis and conclude that there is a
Blend X statistical difference between the blends.
Count 5 5 5 5 20  Since the p-value (crops) = 0.0649 > 0.05 = α, we can’t reject
Sum 659 677 879 706 2921
Average 131,8 135,4 175,8 141,2 146,05
the Factor B null hypothesis. So we can conclude (with 95
Variance 844,2 707,8 278,7 354,2 782,3658 percent confidence) that there are no significant differences
between the effectiveness of the fertilizer for the different
Blend Y crops.
Count 5 5 5 5 20  We also see that the p-value (interactions) = 0.0456 < 0.05 =
Sum 716 798 701 827 3042
α, and so we conclude that there are significant differences in
Average 143,2 159,6 140,2 165,4 152,1
Variance 498,7 978,3 165,7 217,3 511,0421
the interaction between crop and blend.
 Follow-up tests can give more ideas for deepening the
Blend Z analysis.
Count 5 5 5 5 20
Sum 822 868 932 862 3484 ANOVA
Average 164,4 173,6 186,4 172,4 174,2
Source of Variation SS df MS F P-value F crit
Variance 443,3 428,8 212,3 175,8 330,6947
Sample (BLENDS) 8782,9 2 4391,45 9,933347 0,000245 3,190727
Total Columns (CROPS) 3411,65 3 1137,217 2,572355 0,064944 2,798061
Count 15 15 15 15 Interaction 6225,9 6 1037,65 2,347138 0,045555 2,294601
Sum 2197 2343 2512 2395 Within 21220,4 48 442,0917
Average 146,4667 156,2 167,4667 159,6667
Variance 705,8381 871,0286 605,981 404,9524
Total 39640,85 59

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY


Beyond our study: Factor and Cluster Analysis

Factor analysis is based on the idea that certain variables ‘go together’,
in that people with a high score on one variable also tend to have a high
score on certain others, which might then form a group
 Example: People who go to the theatre might also visit galleries

Cluster analysis is another ”grouping” procedure, but it focuses on the


individuals directly rather than the variables.
 Example: Imagine a situation with two variables, age and some behavioural variable,
and data points plotted in the usual way. It can be seen that there are three broad
‘clusters’ of respondents – two young clusters and one older cluster. Each of these
clusters might form, for example, particular market segments.
 With just two variables and a few observations it is relatively simple to identify clusters
visually. But with more variables and more cases this would not be possible. Cluster
analysis involves giving the computer (SPSS or R) a set of rules for building clusters.

Jozsef Zoltan Malik INTRODUCTORY RESEARCH METHODOLOGY

You might also like