0% found this document useful (0 votes)
4 views636 pages

Statistics For The Social Sciences

Uploaded by

wangruosongsong
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views636 pages

Statistics For The Social Sciences

Uploaded by

wangruosongsong
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Digitized by the Internet Archive

In 2022 with funding from


Kahle/Austin Foundation

[Link]
STATISTICS
Ort
sCIA.
SCIENCES
To the memory of my wife, Debbie, and to my daughters, Edie and Farah.

‘All my pretty chickens and their dam.”

And, a new addition for a new edition:

To my grandson, Alexander Duncan Mitchell, born to Edie and her husband, Robert.
STATISTICS -
for the
SOCIAL
SCIENCES
Wee PRED E Dalalali@ain

Thousand Oaks # London #® New Delhi


Copyright © 2006 by Sage Publications, Inc.

All rights reserved. No part of this book may be reproduced or utilized in any form
or by any means, electronic or mechanical, including photocopying, recording, or by
any information storage and retrieval system, without permission in writing from the
publisher.

For information:

Sage Publications, Inc.


§) 2455 Teller Road
Thousand Oaks, California 91320
E-mail: order@[Link]

Sage Publications Ltd.


1 Oliver’s Yard
55 City Road
London EC1Y 1SP
United Kingdom

Sage Publications India Pvt. Ltd.


B-42, Panchsheel Enclave
Post Box 4109
New Delhi 110 017 India

Printed in the United States of America

Library of Congress Cataloging-in-Publication Data

Sirkin, R. Mark.
Statistics for the social sciences / R. Mark Sirkin.—3rd ed.
p. cm.
Includes bibliographical references and index.
ISBN 1-4129-0546-X (pbk.)
1. Social sciences—Statistical methods. 2. Statistics. I. Title.
HA29.S5763 2006
519.5—dce22
2005007296

This book is printed on acid-free paper.

OB OG 107 As Oe ea GSR BI

Acquiring Editor: Lisa Cuevas Shaw


Editorial Assistant: Karen Gia Wong
Production Editor; Diana E, Axelsen
Copy Editor; Gillian Dickens
Typesetter: C&M Digitals (P) Ltd.
Indexer: Judy Hunt
Cover Designer: Janet Foulger
Contents

Preface

Acknowledgments

Note to Students g
E
a

1. How We Reason
Key Concepts
Prologue
Introduction
Setting the Stage
Science
Example
The Scientific Method ee
ese
Re
RO
SRG
NG,
Testing Hypotheses
From Hypotheses to Theories
Types of Relationships
Association and Causation
The Unit of Analysis
Example
Conclusion
Exercises ND
DN
NMGy
CO
tS
Ul
GN
9=
©
WON
HB
He

2. Levels of Measurement and Forms of Data QnQ

Key Concepts
Prologue
Introduction
Measurement
Qualitative and Quantitative Data
Nominal Level of Measurement CyOb
Ov
OoOR
AR
WV
GN
©
Ovo
Oe
Cv

Forms of Nominal-Level Data 38


Ordinal Level of Measurement
_ Forms of Ordinal-Level Data
Likert Scales
Scores Versus Frequencies
Interval and Ratio Levels of Measurement 45
Forms of Interval-Level Data 48
Tables Containing Nominal Level of Measurement Variables 54
Conclusion 55
Exercises 56
3. Defining Variables 63
Key Concepts 62
Prologue 63
Introduction 64
Gathering the Data 64
Operational Definitions 65
Index and Scale Construction : 69
Validity Bie,
Reliability JE)
Conclusion a
Exercises 80

4. Measuring Central Tendency 83


Key Concepts 82
Prologue 83
Introduction 84
Central Tendency 84
The Mean 85
The Median 90
Grouped Data 94
Using Central Tendency 98
The Mode 98
Interpreting Graphs 104
Central Tendency and Levels of Measurement 105
Skewness 106
Other Graphic Representations 112
Stem and Leaf Displays i
Boxplots 113
Conclusion 116
Summary of Major Formulas iy
Exercises 1

5. Measuring Dispersion 127


Key Concepts 126
Prologue hey
Introduction 128
Visualizing Dispersion 128
The Range 129
The Mean Deviation 130
The Variance and Standard Deviation 32
The Computational Formulas for Variance
and Standard Deviation 136
Variance and Standard Deviation for
Data in Frequency Distributions 139
Conclusion 141
Summary of Major Formulas 141
Exercises 142

6. Constructing and Interpreting Contingency Tables 149


Key Concepts 148
Prologue 149
Introduction 150
Contingency Tables 150
Regrouping Variables 15
Generating Percentages 155
Interpreting 159
Example 160
Controlling for a Third Variable 164
Partial Tables 167
Causal Models 72
Computer Applications iS
oy ee 175
SAS 177
Conclusion 180
Exercises 181

7. Statistical Inference and Tests of Significance 191


Key Concepts 190
Prologue teal
Introduction 192
What Is Statistical Inference? 192
Random Samples 195
Comparing Means 198
Comparing a Sample Mean to a
Population Mean or Other Value 200
Comparing a Sample Mean to
Another Sample Mean 202
Comparing More Than Two Sample Means 202
The Test Statistic 203
Probabilities
Decision Making
Review
Examples
Directional Versus Nondirectional Alternative
Hypotheses (One-Tailed Versus Two-Tailed Tests)
Setting the Level of Significance
Degrees of Freedom
Conclusion
Steps in Significance Testing
Summary of Major Formulas
Exercises

8. Probability Distributions
and One-Sample z and t Tests
Key Concepts
Prologue
Introduction
Normal Distributions NNN
The One-Sample z Test for Statistical Significance WN
DNns
ron
The Central Limit Theorem
Review - Ov

The Normality Assumption


The One-Sample ¢ Test LY
DNDN
wNN

Degrees of Freedom
The ¢ Table
An Alternative ¢ Formula
A z Test for Proportions
Interval Estimation
Confidence Intervals for Proportions
More on Probability NWN
LH
WK DVWN
©
R
On
NNONNN
BS
Sv
WV

The Addition Rule


The Multiplication Rule
Permutations and Combinations
Conclusion
Summary of Major Formulas
Exercises

9. Two-Sample ¢ Tests
Key Concepts
Prologue
Introduction De
Independent Samples Versus Dependent Samples gue
The Two-Sample ¢ Test for
Independently Drawn Samples Pati)
Adjustments for Sigma-Hat Squared (6°) 288
Interpreting a Computer-Generated ¢ Test 288
Computer Applications: Independent Samples ¢ Tests 290
SPSS: 290
SAS 221
Excel 294
The Two-Sample ¢ Test for Dependent Samples 297,
Computer Applications: Dependent Samples ¢ Test 301
SESS 301
SAS 301
Excel 303
Statistical Significance Versus Research Significance 303
Statistical Power 306
Conclusion 309
Summary of Major Formulas 309
Exercises oun
10. One-Way Analysis of Variance 317
Key Concepts 316
Prologue Sly
Introduction 318
How Analysis of Variance Is Used 318
Analysis of Variance in Experimental Situations ayy)
F: An Intuitive Approach 522
ANOVA Terminology 326
The ANOVA Procedure 330
Comparing F With ¢ oF
Analysis of Variance With Experimental Data 338
Post Hoc Testing 340
Computer Applications 343
SESS. 343
SAS 345
Excel 348
Two-Way Analysis of Variance 348
Conclusion 552
Summary of Major Formulas 552
Exercises 559
11. Measuring Association in Contingency Tables 359
Key Concepts 358
Prologue 359
Introduction 360
Measures for Two-by-Two Tables 360
Yule’s Q 362
The Phi Coefficient 305
Measures for 7-by-72 Tables 367
Goodman and Kruskal’s Gamma (Yy) 367
Goodman and Kruskal’s Lambda (i) a7
Lambda—Column Variable Dependent ale
Lambda—Row Variable Dependent ae,
Curvilinearity ore
Other Measures of Association 380
Interpreting an Association Matrix 381
Conclusion 385
Summary of Major Formulas 385
Exercises 386
12. The Chi-Square Test 397
Key Concepts 396
Prologue 397
Introduction 398
The Context for the Chi-Square Test 398
Expected Frequencies 400
Observed Versus Expected Frequencies 405
Using the Table of Critical Values of Chi-Square 408
Calculating the Chi-Square Value 412
Yates’s Correction 415
Validity of Chi-Square 417
Directional Alternative Hypotheses 422
Testing Significance of Association Measures 425
Association Versus Significance 426
Chi-Square and Phi 429
Computer Applications 431
SPSS. 431
SAS 433
Conclusion 435
The Limits of Statistical Significance 435
Summary of Major Formulas 436
Exercises 437
13. Correlation and Regression Analysis
443
Key Concepts
442
Prologue
443
Introduction 444
The Setting 444
Cartesian Coordinates 447
The Concept of Linearity 451
Linear Equations 455
Linear Regression 460
The Correlation Coefficient 408
The Coefficient of Determination 472
Finding the Regression Equation 474
Computer Applications 479
SPSS 479
SAS 484
Excel 485
Correlation Measures for Analysis of Variance 486
Conclusion 489
Summary of Major Formulas 489
Exercises 490
14. Additional Aspects of
Correlation and Regression Analysis 497
Key Concepts 496
Prologue 497
Introduction 498
Statistical Significance for r and b 498
Significance of r 506
Partial Correlations and Causal Models 508
The Role of the Partial Correlation Coefficient si
Multiple Correlation and the
Coefficient of Multiple Determination 516
Multiple Regression 520
An Example From Judicial Behavior 524
The Standardized Partial Regression Slope 528
Using a Regression Printout 530
Stepwise Multiple Regression B65)
Computer Applications 538
Partial Correlations—SPSS 538
Partial Correlations—Other Programs 540
Multiple Regression—SPSS 540
Multiple Regression—SAS 543
Multiple Regression—Excel 547
Stepwise Multiple kegression—SPSS 547
Stepwise Multiple Regression—SAS 552
Conclusion 552
Summary of Major Formulas 57
Exercises 558

Appendix 1: Proportions of Area


Under Standard Normal Curve 561
Appendix 2: Distribution of t 565
Appendix 3: Critical Values of F for p = .05, .01, and .001 566

Appendix 4: Critical Values of Chi-Square 569


Appendix 5: Critical Values of the Correlation Coefficient 570
Answers to Selected Exercises S71
Glossary 587

Index 603
About the Author 610
Preface

o the students and the instructors using this text, welcome!


This book is designed to teach introductory statistics primarily
to undergraduates majoring in the social sciences. I have tried
to use a wide variety of examples that are both relevant to the social and
behavioral sciences and of interest to today’s undergraduate students. This
book may be used as a text in a statistics course geared to any of the social
sciences, or it may be used as part of a course or sequence of courses in
research methodology.
Why another statistics text? After years of teaching students, many of
whom claim to be victims of math anxiety, I wanted to provide a teaching
device that could be used by the non-mathematically inclined, but at the same
time would cover all relevant topics thoroughly enough to meet the needs of
all students. To do this, (a) Iam assuming that the only recent math courses
readers have had did not go much beyond introductory algebra and (b) many
of the more onerous calculations encountered can be done on a computer.
So, while all relevant calculations are presented here, emphasis—particularly
in the later chapters—is also placed on the submission of computer runs and
the analysis of the computer outputs generated from those runs.
Another thing that I have done is to begin with as little computational
work as possible and move slowly into the math. This approach should
enable students to gradually overcome their fear of numbers, build confi-
dence in their ability to handle quantitative work, and (who knows?) even
come to enjoy what they are doing. Note that many of the earlier topics,
such as those on the scientific method, levels of measurement, and inter-
pretation of tables, are given far less attention in many other Statistics texts
than is given here. By including them, it is my hope that students will see
statistics as linked to the more comprehensive field of research methodol-
ogy, rather than just as an entity unto itself. My emphasis is on the analysis
and interpretation of data, rather than on how those data are collected.
However, I do want the reader to have a feel for the way interpretation of
data is related to the methods whereby the data were obtained. This
approach also guards the student from immediate inundation in calculation.

> xiii
xiv STATISTICS FOR THE SOCIAL SCIENCES

Examples and exercises are designed to mirror the subject matter


explored in all the social science disciplines. Easily spotted throughout the
text are examples that can be identified with sociology, political science,
communications, psychology, social work, management, education, and
other disciplines. In selecting examples, I have chosen topics that should be
of general interest to undergraduates in each field or to all undergraduates
in the social sciences. I have also sought to include examples that reflect
applied research as well as basic research. Such examples and exercises
should help students retain interest in the course material. This has been my
experience with my own students after having “test marketed” drafts of
these chapters.
Each chapter begins with a prologue, an introduction, and a list of key
concepts that are introduced for the first time in that particular chapter. In
the body of each chapter, key concepts are presented in boldface type with
their definitions separated to stand out. Boxes are used to provide supple-
mental information reinforcing certain topics presented in the chapter.
Where appropriate, mathematical formulas are summarized at the end of
the chapter. The exercises at the end of the chapter are presented in the
same order in which the material is presented within the chapters, so that
they may be undertaken prior to the completion of all topics presented in
the chapter. There are ample exercises so that instructors may assign at their
discretion a subset of the problems and still cover all the appropriate statis-
tical procedures presented in the chapter. Thus more homework problems
are included than one may need to assign. Some additional homework
problems have been added to the third edition; in many cases, these reflect
the additional computer coverage added to this book’s contents as well as a
revision of the data used in Chapters 6, 11, and 12.
To the instructors, a word about the sequence of chapters is in order.
Perhaps no other issue has taken as much thought and discussion time as
this. There are many ways of presenting those topics, and all instructors
have personal preferences that are equally valid but difficult to reconcile
with one another. The ordering I have used is a compromise between dif-
fering approaches. The rationale for this ordering will be addressed, but
prior to doing so, let me emphasize that as far as possible, the chapters are
written in such a way that they may be presented to students in alternative
ordering patterns without great difficulty. After reading the following para-
graphs, you should be able to easily present the material according to your
own needs and preferences.
Chapters 1 through 6 are designed to introduce students to concepts
of empirical research and the basic working vocabulary of statistics. These
chapters cover the scientific method, levels of measurement and formats
Preface ® XV

for manipulating and presenting data, operational definitions and index


construction, central tendency and dispersion measures, and contingency
tables. Although all introductions to statistics cover central tendency and
dispersion, few give as much emphasis to the other topics as I have given in
this book. This additional coverage will be of particular value if you are using
this text in a combined methods/statistics course or if students have not
already taken a separate methods course.
Chapters 7 through 10, together with Chapter 12, cover inferential
statistics. Chapter 7 is an overview of the entire area and should be read
first. From that point, it is possible to cover the remaining inferential statis-
tics chapters in any order, except that Chapter 11 must be read prior to
Chapter 12. Although generally one presents the two-sample ¢ test prior to
covering analysis of variance, it is possible, for instance, to cover chi-square
without having first covered ANOVA.
Chapter 11 (association measures) and Chapter 13 (linear regression)
cover additional topics in descriptive statistics and could be presented prior
to the chapters on statistical inference. Chapter 14, however, can be fully
utilized only if students have already had several of the inferential statistics
topics in addition to linear regression. If you have a two-course sequence, it
is possible to group the descriptive statistics chapters in the first course
(Chapters 1-6, 11, and 13) and the inferential chapters in the second course
(Chapters 7-10 and 12), culminating with Chapter 14, which interweaves
the two threads. In short, I have designed the chapters with the knowledge
that there are many possible sequences of topics and that we all march to
difference drummers.
In this edition, the computer runs have all been updated to their most
recent versions, and coverage has been extended to include Excel.
For the instructor, I will save more of my comments for the Instructor’s
Resources CD that accompanies the book. The CD has been expanded with
considerably more material than the earlier manual, including a list of test
questions and other material. The CD also includes suggestions gleaned
from years of teaching this material. In turn, I hope that you will share your
observations—both positive and negative—with me. You may contact me in
care of the Department of Political Science, Wright State University, Dayton,
OH 45435. Best wishes for a positive teaching and learning experience.
. it. re _ rr ae ree > as

Aah P7a~en ete Beiets a


= Saas an, Mee =o o's a» See
ha aid 7 ey Or fast eH pe
eaten aoe
Acknowledgments

here are many people who contributed directly or indirectly


to this book. Indirectly, all those professors and colleagues
who helped me develop my interest in statistics and method-
ology deserve my thanks. Likewise, my family deserves my gratitude
for their patience and support during a period when the demands of the
book interfered with my family activities and responsibilities. Finally, |owe
a debt to all of my colleagues in the Wright State University Department of
Political Science and the College of Liberal Arts for their encouragement
and support.
Of those who contributed directly, nobody deserves greater praise than
Joanne Ballmann, who typed, retyped, and often re-retyped all editions of
the manuscripts. I don’t think that Joanne realized how difficult a job it
would be to prepare a text so full of symbols, Greek letters, and algebraic
equations. However, she tackled the project with much patience and good
humor. Thanks! My gratitude also goes to Bruce Stiver and Gloria Sparks, in
the graphic arts office of Wright State’s Media Services, who prepared the
many figures and graphs found throughout this book.
Thanks also to the College of Liberal Arts at Wright State and its dean,
Mary Ellen Mazey, and its former dean, Perry Moore, now Provost and Vice
President for Academic Affairs at Texas State University, San Marcos. The
college provided me with funds for travel, graphics, and manuscript pre-
paration. In the process of gaining this seed money, I received support and
assistance also from Jim Jacob, now at California State University, Chico;
Charlie Funderburk, my former department chair; and Donna Schlagheck,
my current chair. Also, thanks to Bill Rickert, currently Associate Provost,
and, of course, my colleagues on the college’s Faculty Development
Committee. And to my many friends at Wright State who provided me with
ideas and examples from their various social science disciplines, my grati-
tude goes out to you.
Many faculty members at universities and other scholars read and
commented on drafts of chapters for the first edition. Specifically, I would
like to thank professors Rick Brown, California State University at Fresno;

od XVil
XVill @ STATISTICS FOR THE SOCIAL SCIENCES

Alfred DeMaris, Bowling Green State University, Ohio; David Dooley,


University of California at Irvine; Donald Gross, University of Kentucky,
Lexington; Carl J. Huberty, University of Georgia, Athens; Garth Lipps,
Statistics Canada, Ottawa; and John P Mclver, University of Colorado,
Boulder. Many helpful reviewers provided useful assistance during the
course of the second edition of this manuscript’s development: David
Dooley, University of California, Irvine; Larry Marsh, University of Notre
Dame; Jerome McKean, Ball State University; Steve Seitz, University of
Illinois; Bruce A. Thyer, Ph.D., University of Georgia; Holly Gimpel, St.
Joseph’s College; Bryn Geer-Wootten, York University; and Mirka Ondrack,
York University.
Thanks to the following reviewers for the third edition: Christopher
Hiryak, Arizona State University, School of Public Affairs; Amilcar Antonio
Barreto, Northeastern University; Robert Mark Silverman, Department of
Urban and Regional Planning; State University of New York at Buffalo;
Dennis W. Roncek, University of Nebraska at Omaha; Lesley Andres,
University of British Columbia; and Catalina Stefanescu, London Business
School.
For the first two editions, thanks to the following people at Sage
Publications: C. Deborah Laughton, Senior Editor, and my production
editors Diane Foster, who edited the first edition, and Diana Axelsen, who
edited the second and third editions. For the third edition, my thanks to
Lisa Cuevas-Shaw, Acquisitions Editor, and in addition to Diana Axelsen,
Karen Wong, Margo Crouppen, Gillian Dickens, and A. J. Sobczak. Thanks
also to editorial assistants Nancy Hale (first edition), Eileen Carr (second
edition), and Karen Wong (third edition), who were my main day-to-day
contacts in California. Acknowledgments also to Tricia Howell Bennett
and Andrea Swanson, who assisted with the first edition, and Christina
Hill, who was typesetter for the first edition and typesetting coordinator
for the second edition. Let me also thank Stephanie Caballero, Promotions
Manager, for her efforts. Many others at Sage deserve my thanks as well.
Finally, my gratitude to the people at Technical Typesetting, Inc., Baltimore,
Maryland, for their work on the second edition, and to the staff of C&M
Digitals (P) Ltd. for making the third edition as flawless and professionally
pleasing as possible. Iam indebted to all of you!
Also, one learns from one’s students. I wish to thank my students at
Wright State University who took my classes in quantitative methods during
the past years and used earlier drafts of the chapters as text material. Their
comments and feedback contributed greatly to improvements I was able
to make in these chapters. In particular, Marge Gibson, Ann Koch, and
Connie Weber, three of my students, were kind enough to supply detailed
Acknowledgments >» xix

commentary—and proofreading—for several draft chapters. Thanks also


to Chang Li for assisting in the index preparation and to Farah Sirkin for
helping type the index for earlier editions.
From Wright State University, special thanks to Mike Kepler, Computing
and Telecommunications Services, and Virginia Gimenez, College of Liberal
Arts, for great assistance. with the hardware and software used to produce
this edition.
For sources cited in this book, my thanks to the following:

Americans for Democratic Action for permission to use their ratings


of U.S. Representatives, as reproduced by CQ Press in H. Stanley and
R. Niemi (1992), Vital Statistics in American Politics (rd ed.)

Freedom House ([Link]) for use of data from its Web


site, Freedom in the World, 2003

McGraw-Hill/Dushkin Publishing, Dubuque, IA, for use of data


summarized in William Spencer (2000), Global Studies: The Middle
East (8th ed.)

PearsonEd EMA (Pearson Education, London), on behalf of the Literary


Executor of the late Sir Ronald A. Fisher, F.R.S., and Dr. Frank Yates,
F.R.S., for permission to reproduce Tables II-1, I, IV V and VII from
Statistical Tables for Biological, Agricultural and Medical Research
6/e (1974)

Profile Books, Ltd., London, to publish data from The Economist: Pocket
World in Figures, 2004 edition
‘Transparency International ([Link]/surveys/index,html#cpi)
for use of its Corruption Perceptions Index 2002 data

Yale University Press permitted my use of data and excerpts from


C. L. Taylor and D. Jodice (1983), World Handbook of Political and Social
Indicators. Although almost all of those data are mot used in this edition,
they are referenced and in one or two instances reproduced here.
Finally, my thanks to the companies whose software was used in the
computer applications presented in this edition and the accompanying
Instructor’s Manual: the SAS Institute for use of SAS Version 8.2, SPSS for
use of SPSS 12.0 for Windows, and Microsoft for use of the Excel statistical
routines.
Others too numerous to mention have contributed to this textbook.
To all of them, thanks! Any errors to be found are, of course, not theirs
but mine.
xx @ STATISTICS FOR THE SOCIAL SCIENCES

Created with SAS® software. Copyright © 2004, SAS Institute Inc., Cary,
NC, USA. All Rights Reserved. Reproduced with permission of SAS
Institute Inc., Cary, North Carolina.
SPSS® is the registered trademark of SPSS Inc., 223 South Wacker
Avenue, Chicago, Illinois, 60606-6307. SPSS program code, output,
and execution logs are reprinted with the permission of SPSS Inc.
SPSS 12.0: Copyright © 2004, SPSS Inc., Chicago, IL. All rights reserved.
Reprinted with permission.
Microsoft Excel®: Microsoft product screen shots reprinted with
permission from Microsoft Corporation. Excel® is the registered trade-
mark of Microsoft Corporation, One Microsoft Way, Redmond, WA
98052-6399. Reproduced with permission of Microsoft Corporation,
Redmond, WA.
Note to Students

nlike many other courses, the material presented here is often


cumulative in nature. This means that to understand today’s
assignment, you need to understand the material previously
presented. If you do not understand today’s topic, you may not be able
to understand tomorrow’s. Accordingly, do not try to “cram” this material.
Learn it at a regular pace. If something confuses you, stop and reread it.
If you still do not understand the topic, ask your instructor. Don’t feel self-
conscious about raising your hand in class and asking. You need to under-
stand the material! Moreover, if you are confused, the odds are that others
are likewise having trouble. Never assume that you are the only one
“snowed.” One other suggestion: Attend class! You will learn better with the
reinforcement provided by your instructor in class, and you will have your
instructor at hand to help explain any material that is causing you difficulty.
Statistics calls for class attendance.
It is my hope that this book will help contribute to a worthwhile educa-
tional experience for you. Best wishes with it!

Xxi
VY KEY CONCEPTS ¥
LINN PORE

empirical concept positively related/


normative variable a positive relationship
scientists induction off diagonal
hypothesis/ deduction inversely related/an
hypotheses experiment inverse relationship
social sciences scientific law causation
scientific method data (pl.)/datum temporal sequence
table/cross-tabulation/ or piece dependent variable
contingency table of data (sing.) independent variable
marginal totals necessary condition criterion variable
grand total sufficient condition predictor variable
cell (ofa table) theory unit Of analysis
association main diagonal Statistics
NR
CHAPTER

How We Reason

VY PROLOGUE ¥
BREESE
ESS SITES SELLE NILE
IO EI SIE LN LN DALELEO NER ELEN DOELELEL LEE ESEDNL AEES OEE SEVERE ALS SSBOLI T OSES EE LE LEER [Link]

Sometimes I wish I were not a political scientist. Unlike some other social
sciences, political science also has deep roots in the humanities and in
nonquantitative research, and this has led to countless debates in my field
concerning epistemology, the branch of philosophy dealing with what is
knowledge, and how we should study politics. In a nutshell, the debate
is over whether quantitative methods and statistical techniques provide a
better picture of reality than more traditional, nonquantitative scholarship.
I have always been a generalist who takes knowledge any way it is
available, without dealing with the nature of how we know if what we see
actually is what is. Like the late U.S. Supreme Court Justice Potter Stewart,
who in 1964 commented that although he couldn’t define pornography, he
knew it when he saw it, I sometimes wished that the debate would go away
and, like psychologists or physicists, we could just “do science,” without
having to justify what we do to doubting colleagues.
But, on the other hand, maybe there is a role for epistemology. What do
we mean when we “do science”? What is science anyway, and how do the
social sciences relate to science in general? Before we can “do science,” we
must know what science is.
BSERSS eae ate RITA I AG RAENAEEIS SRE 2 ENE ORT E S R REII OOS
2 << STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
This is a textbook on statistics and data analysis for the social sciences. Its
techniques apply whenever data that involve counting or measuring have
been collected. A standard logical process—the scientific method—underlies
the collection and interpretation of data and applies to all the sciences.
The scientific method is the procedure whereby we propose possible
relationships among characteristics of phenomena under study and then
test to see whether those relationships actually exist. Although the tech-
niques for doing this vary from one discipline to another, the logical
sequence remains the same. Therefore, the scientific method provides a
good jumping-off point for the chapters to come. This, in other words, is
how we reason.

SETTING THE STAGE


All social sciences—indeed, all sciences in general—have many goals.
Essentially, we seek to understand some relevant phenomenon to make pre-
dictions or provide explanations about that phenomenon and possibly to
gain some control over it. To do this, we generally find ourselves concerned
with two fundamental kinds of questions—What is? and What ought to be?
Questions pertaining to knowing “what is” we call empirical, and questions
about “what ought to be” we call normative. These two types of questions
are found side by side in all academic disciplines, with the normative pre-
dominating in the humanities, the empirical predominating in the labora-
tory sciences, and both being components of the social sciences.

Empirical questions Questions that pertain to knowing “what is.”

Normative questions Questions that pertain to “what ought to be.”

Questions about what ought to be form a core essential to understand-


ing our society. We ask such questions as these: What is a good society? What
is justice? Is a democratic government an ideal form of political system? Is it
moral for opinion leaders to use deceit in dealing with the public? Should
we employ the death penalty in the case of premeditated murder? Is racial
discrimination morally acceptable? The answers to these questions depend
on our values, Our priorities and preferences, and our feelings of what is
right or just. When we ask, “What are the components of an ideal society?”
we are really asking ourselves to specify what we as individuals desire
or think contributes to the common good. Normative questions form the
How We Reason ® 3

background of social and political philosophy. They extend as far back as


human preference can be traced, and some of the greatest minds have
sought their answers: Plato, John Locke, John Stuart Mill, Thomas Jefferson,
Adam Smith, and others—the giants of normative theory.
The empirical questions pertain to establishment of facts rather than
values. We may accept the normative ideal that freedom of the press is a
necessity for a free society. The related empirical question would be, Is there
freedom of the press in a given nation? and/or How much freedom of the
press actually exists there? Suppose we define freedom of the press as the
existence of at least one newspaper or TV station not owned or controlled
by the government and not subject to censorship by the government.
Whatever we may like or dislike about that definition of freedom of the
press, once we tentatively agree to use it as our working definition, we are
in a position to examine data about each country and, in doing so (perhaps
with the aid of experts), to determine whether freedom of the press really
exists there. The existence of freedom of the press as we defined it is subject
to empirical verification. The question of whether freedom of the press is
good and should exist in a democracy is a normative question, a matter of
personal preference.
Both normative and empirical considerations enter into any definition
of freedom of the press. For instance, we have many examples where
countries with wide peacetime press freedoms must impose military cen-
sorship in times of war for security purposes. Our values will determine
whether we would still categorize that country as having a free press. What
about restrictions resulting from libel laws? What about nudity? What kinds
of restrictions may we accept and still consider the press to be free?
Empirical considerations pertain to our need to be able to clearly measure
what we study. How are we going to obtain the information needed to cate-
gorize a country’s press freedom? If a free press, in our definition, allows
“soft-core” but not “hard-core” pornography to be published, how are we to
determine where one ends and the other begins?
The normative and the empirical coexist and complement one another in
a simple way: Empirical facts tell us how close we are to the normative ideal
or how far we still must go to achieve the ideal. Suppose we have normative
agreement that poverty should be eliminated. If we know that 10 years ago,
15% of all families were legally defined as living below the poverty level and
this year only 12% can be so defined, we might conclude, at least initially, that
poverty is declining and we are moving toward our normative goals.
The art of empirical analysis—how we establish facts, how we determine
what is, what actually does exist—is the subject matter of this text. Although
many techniques are available for studying “what is,” an underlying logical
process exists to enable us to make use of these analytical techniques.
4 @ STATISTICS FOR THE SOCIAL SCIENCES

EXERCISE
For each of the following, indicate which is normative and which is empirical:

. Abortion should be outlawed.


. There were 3 abortions per 100 live births in the United States last year.
. The right to vote is essential to a free society.
. Only 40% of all registered voters cast votes in the last presidential election.
. Children should not be physically abused.
Last year, the rate of reported child abuse cases rose 3%.
. Aperson who is old enough to fight for his or her country and vote in elec-
tions should be able to purchase liquor legally.
In this state, the minimum age to purchase hard liquor is 21, although the
voting age and minimum age for military service is 18.

Obviously, the odd-numbered questions are normative, and their even-numbered


counterparts are empirical.

A final observation: That something is normative and therefore exists as


a goal rather than an established fact does not make it either more or less
valid to scholars than an empirical fact. Both have a place in the way we
learn about society. Consider the following:

We hold these truths to be self-evident: That all men are created


equal; that they are endowed by their creator with certain unalien-
able rights; that among these are life, liberty, and the pursuit of
happiness... .

These are Thomas Jefferson’s words from the U.S. Declaration. of


Independence. “Self-evident” or not, what do we mean by men (persons)
being created equal? Equal in size? Intelligence? Strength? Equal before the
law? Economically? What about slaves? (Jefferson owned them.) The term
unalienable means that which cannot be taken away. But your right to life
may be taken away by the electric chair, and your right to liberty may be
taken away by a prison term (or a required college course).
Does this mean that what is normative is nonsense? Only if we ignore its
value in providing guidelines and goals for what ought to be or if we ignore
its value in inspiring and motivating us.
How We Reason B® 5

By the same token, the following empirical statement may be true


factually but hardly inspirational.

A recent survey showed that four out of five consumers prefer


Cleeno detergent to the next leading brand.

Unless I were the owner of the company making Cleeno, I would probably
be more inspired by the Declaration of Independence than by the detergent
statement.

SCIENCE

When we focus our attention on empirical knowledge, we concern our-


selves with knowledge based on observation and experimentation. To the
extent that we are engaged in discovering and categorizing such empiri-
cal knowledge, we are being scientists. Scientists are those engaged in
collecting and interpreting empirical information. They do so to formulate
and test hypotheses. Hypotheses, as we use them here, are statements
positing possible relationships or associations among the phenomena being
studied. These relationships, which will be elaborated on throughout this
chapter, suggest that when some attribute or quantity of one phenomenon
exists, a specific attribute or quantity of another phenomenon is also likely
to occur. Because communications, cultural anthropology, political science,
psychology, social work, and sociology, among other disciplines, concern
themselves with aspects of societies, they are often termed the social
sciences. They empirically study social phenomena. Their goals are to
formulate and test hypotheses or suppositions about relationships and
possible causes and effects among various aspects of a society, a culture, or
a political system.

cS HY CR REE SESE LOBIBN SESE STE EOTIERLESSSNS EEIE ESE E ENENER SEES LMS
SEES DEIR OTE SEEING

Scientists People who engage in collecting and interpreting empirical information.

Hypotheses Statements positing possible relationships or associations among the


phenomena being studied.

Social sciences The empirical study of social phenomena.


cwrvarsctseets LLNS TLRSOLSON
IOSE EERE EEEERNE BEES NEESER
LEE IDNREE EN DNS EEL ALLER SALE LNAI ALLS AE NLD NG ROIS

Social sciences share with all sciences two common aspects. The first is
a commitment to the scientific method, a series of logical steps that, if
6 @ STATISTICS FOR THE SOCIAL SCIENCES

followed, help minimize any distortion of facts stemming from the


researcher’s personal values and beliefs. The second is the use of quantita-
tive techniques, measuring and counting, for the gathering and analysis of
the factual information that is collected.

Scientific method A series of logical steps that, if followed, help minimize any
distortion of facts stemming from the researcher's personal values and beliefs.

The scientific method is really a series of intellectual steps. It is not so


much the actual techniques whereby the research is performed as it is the
thought process whereby hypotheses are formed, tested, and verified (or
not verified). If followed, the scientific method provides a basis for acquir-
ing knowledge that eventually will be accepted by the scientific community.
This accepted “truth” would be independent of the values and preferences
of the researcher or any other observer. A properly conducted study of
attitudes on abortion might find that in a particular group—your class, for
instance—55% of the students are pro-choice and 45% are pro-life. This
result would be accepted as fact regardless of the researcher’s personal pref-
erence on the abortion issue.
As we will see, the scientific method involves the formulation of
hypotheses, the testing of hypotheses via observation or experimentation,
and the ultimate verification (or disconfirmation) of the hypotheses in a
manner that will enable other scholars to draw the same conclusion as did
the investigator. Thus, facts are separated from values.
A major challenge for both the researcher and the consumer of research
is to distinguish between what is established fact and what is a moral or eth-
ical value judgment about some aspect of human behavior. This is often not
an easy task because both the selection of facts and the interpretation of the
selected facts filter through the researcher’s own normative screen of values
and preferences.
Among scholars, there has long been a philosophical debate over
whether a science can truly be value free. (This is a question posited against
the claims of those empiricists who desire a value-free science.) Although it
is probably true that we can never totally eliminate the effects of value and
preference, adherence to the scientific method certainly helps minimize
these effects and helps keep those effects from clouding our conclusions.
For instance, a medical researcher convinced that cigarette smoke harms
nonsmokers may publish a report citing the evidence from studies showing
such harmful effects but ignoring those studies showing no damaging
effects. Despite the intervention of personal values in this example, further
empirical studies of passive smoking will either verify existing dangerous
How We Reason ®» 7

effects or show that these effects do not exist. Eventually, the issue
will be
put to rest, despite the occasional encumbrance of personal preferen
ce.
“Facts are facts.”
To enable us to take a closer look at the scientific method, I have chosen
a fundamentally simple example. The example starts with an observation
and proceeds to the establishment of theory.
It should be understood that from theory, further observations are
generated, leading in time to more elaborate theory. Thus, the process is a
cyclical one: observation to theory to observation. Few scientific studies
begin without at least some theoretical foundation. In this example, how-
ever, we assume no prior theory exists.
Now join me in a literal walk through the scientific method, but bring an
umbrella—it might rain.

Example

Suppose each morning I take an hour’s walk. On Sunday, the sun shown
during my walk and I admired the blue, cloudless sky. The same was also true
on Monday. On Tuesday and again on Wednesday, it was raining and the skies
were gray and overcast. On Thursday, it only sprinkled, and the sky, while
generally blue, was broken here and there with a dark cloud. Assuming I lack
education and prior awareness, I note something obvious to you: Rainfall
appears to be associated with the presence of clouds in the sky, whereas sun-
shine generally means fewer clouds. In any event, whenever it did rain, there
were inevitably clouds in the sky. I conclude that rainfall is associated with
cloudy conditions. At least that was the case for this 5-day period.
I think to myself that, if for these 5 days, rain is associated with cloudiness,
then that same pattern should persist over a longer period. I decide to keep
a log, indicating for each day whether I considered it cloudy, partly cloudy, or
clear. I also note for each of these days whether or not it rained. At the end of
30 days, I take stock of my results. I note that of the 30 days, 10 were cloudy,
10 were partly cloudy, and 10 were clear. Also, I note that on 15 days, it
rained, and on the other 15, it did not. I put together a chart, Table 1.1, to
summarize this. We term this chart either a contingency table, a table, or a
cross-tabulation. The totals in the margins of my table are called marginal
totals. The three marginal column totals (10 cloudy, 10 partly cloudy, 10 clear)
add up to a grand total of 30 days. The two marginal row totals (15 rain, 15
no rain) also add up to the same grand total of 30 days.

Contingency table, table, or cross-tabulation A way of presenting data for


purposes of testing hypotheses.
8 ¢ STATISTICS FOR THE SOCIAL SCIENCES

TTT

Marginal totals Row and column totals found in the margins of tables.

Grand total The total number of cases presented in the table. For instance,
in Table 1.1, there are 30 total days being studied.
eR ELEN CCN DN ACO NAA LALLA LL LILIA A LN

Table 1.1

Sky Conditions

Presence of Rainfall Cloudy Partly Cloudy Clear Total

Rain 15
No Rain dhs}
Total 10 10 10 30

Now I review my records and tally my results for the 30-day period, not-
ing for each day what the sky conditions were and whether or not it rained
(see Table 1.2). I count up my tallies and put the appropriate number in
each cell (e.g., “cloudy, rain” or “clear, no rain”) in Table 1.3.

Cell Intersection of a particular row and a particular column. For example,


in Table 1.3, there are five rainy days with partly cloudy sky conditions.

Table 1.2

Sky Conditions

Presence of Rainfall Cloudy Partly Cloudy Clear Total

Rain HN NI III 15
No Rain Ill IAL Unit 15
Total 10 10 10 30

Table 1.3

Sky Conditions

Presence of Rainfall Cloudy Partly Cloudy Clear Total

Rain 10 5 0 15
No Rain 0 5 10 lS
Total 10 10 10 30
How We Reason >» 9

The results for the 30-day period are consistent with the observations
for the initial 5 days: On cloudy days, it always rained; on partly cloudy days,
it sometimes rained; and on clear days, it never rained. The presence of rain-
fall is associated with the presence of clouds, and without clouds, it
appears that no rain will fall. Similar observations over other 30-day periods
of time yield similar results and reinforce my initial conclusions. After a
while, I take the conclusion that clouds are associated with rain for granted.

SSS SSSSSESNOEL AIE ESTE SEESBOSE DEEDS SHS SIERUBSEESSERE ELL BAESSSSSEESLLALLLESE SLL SESS ELE CE EEE EIST AD OE SEEN TEE ES IEEE ENE EERE

Associated When a case falls into a particular category of one concept, such
as rain for the presence of rainfall, it also falls into a particular category of the other,
such as cloudy for sky conditions.
HULL EEESSNS HELLO EELS SEES ONE SE OSES OPES NOORRR HONE ESSENSE NTE

THE SCIENTIFIC METHOD

Aside from my simplistic observational techniques and my simplistic lack of


even a basic understanding of weather, in this little example, I have followed
the scientific method. This is the thought process that is the underpinning
of empirical research. In the first 5 days, I noticed a possible relationship
between two concepts or ideas: sky conditions and the presence of rainfall.
Both of these concepts are known as variables because they vary, or
change, from one observation to another. Sky conditions vary from cloudy
to partly cloudy to clear. Rainfall (as I observed it) varies in the sense that on
some days it rains and on other days it does not. (I could also have classified
my days more specifically if I had wanted to—for instance, heavy rain, mod-
erate rain, light rain, drizzle, no rain. This time I chose to keep it simple.)

se EE SUP LSS SSSI LESSEE RSEEOES ORT PIONS


MORESO LULUIEER BLISS AOM
ULYSSES
SEE LEAL LAE AES

Concepts _ Ideas.

Variables Concepts that vary, or change, from one observation to another.


For instance, some days are rainy; others are not.
EEE SEES NEES LEEDS EENALLE 5 SSSR SSE ESSE ABTS
esa REI SBC TL ESSENSE

The concepts or ideas that we call variables are the phenomena of partic-
ular interest in our social science disciplines. They are called variables because
they vary in amount or attribute for each individual (or group, or society,
or state, or culture—whatever we happen to be observing). Some of the
variables (and their categories or amounts) that we may be trying to better
understand. might include social class (upper, middle, lower), occupational
status
status (white collar, blue collar), political party (Democrat, Republican),
10 « STATISTICS FOR THE SOCIAL SCIENCES

(ascribed, achieved), government (democratic, authoritarian, totalitarian),


or population density (high, medium, low). They parallel the presence of
rainfall during my little walks. Other variables may explain the differences in
the categories of the variables I wish to better understand: income (high,
medium, low—or expressed in actual dollars), education (elementary school,
high school, college—or expressed in total years of schooling), religion
(Catholic, Protestant, Eastern Orthodox, Jewish, etc.), and type of community
(urban, suburban, rural nonfarm, farm). And of course, there may be times
when a variable from the first list, such as social class, might be explaining
variables in the second list, such as income or education.
The relationship that I noticed on my walk was that certain categories of
one variable were associated with specific categories of the other: cloudy sky
conditions with the presence of rain, clear sky conditions with no rain. At
the end of my initial 5 days of observation, I could have stated what I saw as
follows:

There is a relationship between sky conditions and the presence of


rain, such that cloudy sky conditions are associated with the presence
of rainfall and clear sky conditions are associated with no rainfall.

The above statement we term a hypothesis. The hypothesis 7


the two variables that appear to be related and indicates the nature ofthat
relationship (clouds with rain, no clouds with clear weather). Note that
another way of thinking about the relationship in the hypothesis* is an
“if... then” format: If clouds, then there will be rain; if clear, then there will
be no rain. Another but wrong hypothesis could have been that clear skies
are associated with rain; cloudy skies, with no rain. That would have been an
alternative hypothesis, but not one consistent with my initial observations.

Hypothesis A statement that names the variables that appear to be related and
indicates the nature of that relationship.

Here are some examples of hypotheses similar to those one would


expect to find in the social sciences. These examples make use of some of
the concepts presented above:

» There is a relationship between one’s income and one’s social status,


such that the higher one’s social status, the higher will be one’s
income, and the lower one’s status, the lower will be one’s income.
» There is a relationship between the status of one’s occupation
and one’s level of education such that individuals with higher status
How We Reason >» 11

occupations are more likely to have college educations, and individuals


with lower status occupations are more likely to have only grade school
educations,
» There is a relationship in the United States between religious pref-
erence and partisan identity, such that Democrats draw a greater
proportion of suppart from Catholics and Jews than do Republicans,
and Republicans draw a greater proportion of their support from
Protestants than do Democrats.

In each of these cases, despite variations in the wording, the two variables
are named, and the relationship between the categories of each variable is
specified.

Try forming some hypotheses that you think may be true using this same basic
format.

TESTING HYPOTHESES
Based-on a small number of observations during my morning walks (5 days),
I noticed a relationship that I then assumed would hold over one or more
30-day periods. In using a small number of observations to assume that the
relationship should hold for most or all observations, | was undertaking a
process we call induction, going from the specific to the general. I induced
my hypothesis from five specific observations and then assumed that the
hypothesis would apply in all cases.

Induction A process of reasoning that goes from the specific to the general.

Once the hypothesis was induced, I set out to substantiate the


hypothesis by acquiring information over a specific 30-day period. I rea-
soned that if the hypothesis was true in general, it should be true for the
specific 30-day period that I had chosen. Here, I reversed my reasoning and
went not from the specific to the general but the other way, from the gen-
eral to the specific, part of a logical process known as deduction. | deduced
that if the hypothesis were true all the time, it should be true for the 30-day
period I hadselected to study. In other areas of research, I might have been
able to design an experiment to test the hypothesis under laboratory-like
12 << STATISTICS FOR THE SOCIAL SCIENCES

conditions. Here, my “experiment” was to select the 30-day period that I did
and keep records on cloudiness and rainfall.

Deduction A process of reasoning that goes from the general to the specific.

Experiment A test of a hypothesis under laboratory-like conditions.

As the information in Table 1.3 indicates, my hypothesis was verified.


If many subsequent studies had produced similar results, so that the rela-
tionship between presence of clouds and rainfall was widely accepted by the
scholarly community and rarely if ever questioned, then my hypothesis
would become a scientific law.
Scientific laws are simply hypotheses with a high probability of being
correct. There is never absolute certainty about them, despite our use of the
word law. Two facts that emphasize this point will be covered in some detail
later on. First, because most of our data come from sample surveys or
randomized experiments, we are generalizing from a smaller group, the
sample, to a larger one, called a population. However, we never can be
absolutely sure that what was true for our sample is true for the population
as a whole. We only estimate that probability. Second, by the rules of formal
logic, we never in sampling actually prove anything. Instead, we demon-
strate that all other possible alternatives are unlikely to be true, thus leaving
us with only one remaining possibility—the thing that we are proving.

Scientific laws Hypotheses verified so often that they have a high probability of
being correct.

Despite these problems, scientific laws differ from hypotheses in gen-


eral because they are accepted as having a high probability of being correct
by the scholarly community actively pursuing research in the field to which
that scientific law pertains. The law of gravity would be an example. A sci-
entific law is a law because experts in the appropriate area or discipline have
reached that conclusion. It is neither the general public nor scholars in non-
related fields who determine what scientific laws have validity. Physicists, not
sociologists or theologians, determined the validity of gravity.
But the real world rarely cooperates with the researcher as neatly as
indicated in Table 1.3. For example, suppose that after 30 days, I tallied my
results and what appears in Table 1.4 emerged.
How We Reason ® 13

Table 1.4

Sky Conditions

Presence of Rainfall Cloudy Partly Cloudy Clear Total

Rain 5 5 5 15
No Rain oe 5 5 15
Total 10 10 10 30

Table 1.5

Sky Conditions

Presence of Rainfall Cloudy Partly Cloudy Clear Total

Rain 8 8 8 24
No Rain 2 2 2 6
Total 10 10 10 30

Here, the hypothesis is not true; it is disconfirmed rather than con-


firmed. There appears to be 0 relationship between sky conditions and the
presence of rain. Fifty percent of the cloudy days produced rain (5 out of 10
days), but so did 50% of the partly cloudy days and the same percentage of
clear days. Regardless of sky conditions, it rained half of the time. Moreover,
knowing a day’s sky conditions gives us o useful information for predicting
rainfall. Compare this to the information in Table 1.3. There, if we know that
a given day is cloudy, we can predict rain and be 100% correct (all 10 cloudy
days produced rain). If the day is clear, we can perfectly predict no rain (on
no clear day did it rain). The partly cloudy days have rain 50% of the time
only—this is the only category of sky conditions that does not give us per-
fect predictability. If on a partly cloudy day we predict rain, we know that we
can expect to be correct half of the time. But half of the partly cloudy days
produced no rain. We have an equal likelihood of being wrong in predicting
rain. By contrast, in Table 1.4, cloud conditions are not useful at all for pre-
dicting the weather.
At this juncture, note that Table 1.4 indicates a 50-50 chance of rain,
regardless of sky conditions. The 50-50 ratio is the result of the fact that
there are an equal number of rainy and no-rain days, 15 days each, in this
example. One does not need all equal entries to conclude that no relation-
ship exists. Rather, the cell entries need only be proportionate to the mar-
ginal totals. Suppose that out of 30 days studied, it rained 24 days. That
would be 80% of all days studied. If within each category of sky conditions
14. @ STATISTICS FOR THE SOCIAL SCIENCES

it rains 80% of the time, then we also would have no relationship. Assuming
10 days for each of the three weather conditions, that would amount to
8 rainy and 2 no-rain days in each category of sky conditions. This is illus-
trated in Table 1.5.
The information found in these tables we call data. Data is the plural
form. One single piece of information should be called a piece of data or
a datum. Often we forget to differentiate singular from plural, but gram-
matically, we should say “these data” and so on. In Table 1.3, the data con-
firm the hypothesis; in Tables 1.4 and 1.5, they do not.

Data All the information we use to verify a hypothesis.

Datum or a piece of data A single piece of information.

From time to time, a relationship may be found that is 7of in the predicted
direction, as shown in Table 1.6.

Table 1.6

Sky Conditions

Presence of Rainfall Cloudy Partly Cloudy Clear Total

Rain 0 5 10 1S
No Rain 10 ) 0) i)
Total 10 10 10 30

Table 1.7

Sky Conditions

Presence of Rainfall Cloudy Partly Cloudy Clear ~ Total

Rain 8 5 0 13
No Rain 2 5 10 17
Total 10 10 10 30

Here there is a relationship and high predictability, but the relation-


ship does not follow what was logically anticipated by the hypothesis. Clear
days produce rain; cloudy days do not. As in the case of Table 1.4, the origi-
nal hypothesis was not verified, but unlike Table 1.4, there is a relationship
How We Reason iS

between the variables in Table 1.6, only the relationship is the opposite of the
one predicted.
One should note that the relationship found in Table 1.3 is very clear-cut.
Rarely do results appear so clear-cut. More likely it is a case where a trend
is noticeable, even though there are clear examples of days inconsistent with
the hypothesis. Note the illustration in Table 1.7. Only 8 of the 10 cloudy days
resulted in rain; 2 days were inconsistent with the anticipated results.
Nevertheless, the partly cloudy and clear categories remain unaffected. (The
marginal totals for the rows have also changed in this example.) The hypothe-
sis has still been verified, even though the results of the study do not produce
perfect predictability for cloudy days. We may conclude that if the clouds
appear before the rain (clouds come first in time), then cloudy sky conditions
are a mecessary but not sufficient condition for rain. No rain falls without the
presence of clouds, but the presence of clouds does not always result in rain.
We should be aware of this distinction between necessary and sufficient.
A necessary condition is a condition that must be present in order for
some outcome (in this case, rain) to occur. Its presence, however, does not
guarantee that the outcome will occur. By comparison, if a sufficient con-
dition exists, the predicted outcome will definitely take place. For example,
one could argue that poverty is a cause of communist revolutions. Indeed,
the presence of poverty motivated Marx, Lenin, and Mao in their writings
and strategies, and there was great poverty in prerevolutionary Russia and
China. Yet, many impoverished nations have not undergone Marxist revolu-
tions. Why a revolution in Cuba but not in Haiti? Perhaps poverty is neces-
sary but not sufficient for such a revolution. Then, in addition to poverty,
one or more other factors may be needed for a revolution, such as a per-
ception of inequality, unmet rising expectations of an end to poverty, or an
organized revolutionary movement. If the presence of poverty alone always
led to leftist revolution, then it would be both necessary and sufficient. It is
also possible that any of several conditions, when accompanying poverty,
can cause revolution; for example, either poverty plus a perception of
inequality or poverty plus a charismatic revolutionary leader is sufficient to
bring about revolution. When we study hypotheses containing more than
two variables, we take the necessary versus sufficient aspect of relationships
into particular consideration.
il

Necessary condition A condition that must be present in order for some outcome
(in this case, rain) to occur.

Sufficient condition A condition in which the predicted outcome will definitely


take place.
OEE ER
EPSMLBA NETL RE TLL RSL EN SELLE LE ETE IA,
ites sz RN BSS
Eee SESSSSNS OEE
16 @ STATISTICS FOR THE SOCIAL SCIENCES

FROM HYPOTHESES TO THEORIES

Much in the same manner as my rainfall study, I could design studies to


test my social science hypotheses. Perhaps I am interested in the possible
relationship between religion and political party preference. I could pre-
pare an attitude questionnaire and administer it to a randomly selected
group of people. One question would ask each person his or her party
preference, and another question would tap religious preference. From
the results, we could put together a table similar to those in the rainfall
example (see Table 1.8).

Table 1.8

Religious Preference

Party Preference Protestant Catholic Jewish Eastern Orthodox _ Total

Republican
Democrat

Total

We would then look for discrepancies in the proportion of each reli-


gious group expressing preference first for the Republicans and then for the
Democrats.
Let us assume that the findings show what we anticipated—Catholics
and Jews have a greater tendency to express preference for the
Democratic Party than do Protestants, and Protestants have a greater ten-
dency to identify with the Republican Party than do the two other religious
groups. This would motivate me to expand my study. Perhaps I might use
as an explanatory variable family origin (for example, the country from
which the respondents’ forebears came when immigrating to the United
States). I might look at the communities where the immigrants settled and
the freedom and social mobility available to those immigrants. I might
look at the impact of the Great Depression or the civil rights movement
on the voting patterns of the families of my respondents. If Ican establish
a number of interrelated hypotheses, each of which partly explains or
accounts for differing partisan identities, 1 would be developing a theory
of partisan identity. The theory must also account for as many exceptions
to the rule as possible. The better the theory, the fewer exceptions there
will be.
How We Reason >» ie
enn
Theory A set of interrelated hypotheses that together explain some phenomenon
such as why one identifies with a particular political party.
enemies

Into my theory would pour dozens of observations and hypotheses,


some already known and others undiscovered. For example:

1. African Americans were pro-Republican following the Emancipation


Proclamation. As a result, however, of post-Reconstruction Jim Crow
legislation, African Americans in the South abandoned the Republican
Party.

2. Southern Whites, opposed to Lincoln’s policies, identified over-


whelmingly with the Democrats until the “Solid South” began to
crumble in the 1960s due to the civil rights movement and the per-
ceived liberalism of the Kennedy and Johnson administrations.

3. Irish and Italian immigrants (mainly Catholic) as well as Eastern


European Jews immigrated to the United States and settled in large
cities, most of which were controlled by the Democrats. Their off-
spring remained Democrat. By contrast, Catholic and Jewish immi-
grants who settled in Republican-controlled communities (a smaller
number) became Republicans.

Eventually, I would have enough explanatory capability to explain why


more Catholics were Democrats than Republicans and why those Catholics
who were exceptions identified with the GOP I would know why Protestants
had a higher probability of being Republicans than people from the other
religion categories. I would know why African and Jewish Americans were
generally Democrats. Finally, I might be able to predict and explain future
trends, such as an impending change in party affiliation by one of these
groups. In short, I would have a theory of partisan identity.
A similar process could take place in the rainfall example. With the
hypothesis tested, I would be inclined to elaborate on it and expand my
study. I might decide to replace the variable sky conditions by one such
as type of clouds present (e.g., cumulus, cumulonimbus, cirrus, etc.) or
humidity level. Perhaps I would also examine temperature or other atmos-
pheric conditions. If I can establish a number of interrelated hypotheses,
each of which partly explains or accounts for the presence of rainfall,
I would have a theory of rainfall. Going to Table 1.7, Iwould need to develop
hypotheses that would, for instance, explain why those 2 cloudy days pro-
duced no rain. What else besides just clouds needs to be present if rain is to
18 @ STATISTICS FOR THE SOCIAL SCIENCES

fall? Other hypotheses in my theory would explain such exceptions and


more about the nature of rainfall.

TYPES OF RELATIONSHIPS
We worded our original hypothesis this way:

There is a relationship between sky conditions and the presence of


rain, such that cloudy sky conditions are associated with the presence
of rainfall and clear sky conditions are associated with no rainfall.

We named the two variables and went on to specify the nature of the
relationship. When both variables are measured in quantities—as amounts
rather than as differing attributes—it is possible to simplify the specification
of the relationship in our hypothesis. Both variables must be measuring
more or less of an amount. Examples would be net income, either in exact
dollars or categorized as high, medium, or low; age, in years or categorized
as old, middle aged, or young; or liberalism (high, medium, low). Variables
that measure attributes rather than amounts, such as gender, religion,
region, or race, require us to word the hypotheses as we have done so far. To
illustrate such a simplification of our hypothesis, let us recast the categories
of our variables so that both clearly indicate amounts or quantities.
Now in Table 1.9, “Amount of Cloudiness” ranges from most (very
cloudy) to least (not cloudy), and “Amount of Rainfall” ranges from most
rainfall (heavy) to least rainfall (none). The categories of both variables
describe differing amounts, and they are in logical sequence, ordered from
largest to smallest amounts. Note that the vast majority of the 30 days clus-
ter in the table along a diagonal line from upper left to lower right, indicat-
ing that heavier rains are associated with greater amounts of cloudiness and
lighter rains are associated with lesser amounts of cloudiness. Finally, as
shown in Figure 1.1, “no rain” is associated with “no cloudiness.”

Table 1.9

Amount of Cloudiness

Amount of Rainfall Very Cloudy Partly Cloudy Not Cloudy Total

Heavy 7 1 0 8
Moderate 2 4 0 6
Light | 4 0 D)
None 0 1 10 Ld
Total 10 10 10 30
How We Reason 19

Figure 1.1

Amount of Cloudiness

Amount of Rainfall Very Cloudy Partly Cloudy Not Cloudy

Heavy

Moderate

Light
None

We call this diagonal line from upper left to lower right the main
diagonal. In a table with an equal number of rows and columns, the main
diagonal would be a straight line; here it only approximates one. When
most cases cluster on or near the main diagonal, indicating that greater
amounts of one variable are associated with greater amounts of the other
and, conversely, less of one with less of the other, we can describe the
nature of the relationship with the expression positively related. A pos-
itive relationship is one where greater is associated with greater; less
with less.

AeA sees rama See RE SENSE SESE ASSO

Main diagonal Diagonal line from upper left to lower right.

Positive relationship A relationship in which greater is associated with greater;


less with less.
TUE NT i LASSE OL ESE TOON SELES TOE BELT EE DEES IETS,
STILE EEN TUE LEE ELITE

Our hypothesis may now be reworded in simplified form:

The amount of cloudiness and the amount of rainfall are


positively related.

Please note that in each instance, the term positively pertains to the
nature of the actual relationship (high with high, low with low). It is ot
We are
used as a description of how certain we are that a relationship exists.
that the two variables are associated . We are
not saying that we are “positive”
y, positively ” related.
not saying that cloudiness and rainfall are “absolutel
is that the
Rather, we are hypothesizing that the nature of that association
more cloudiness there is, the more rain there will Dey
20 < STATISTICS FOR THE SOCIAL SCIENCES

The opposite of a positive relationship occurs when the variables are


related in a way that the more of one variable there is, the less of the other
there will be. Suppose more clouds mean less rain and fewer clouds mean
more rain, as shown in Figure 1.2.

Figure 1.2

Amount of Cloudiness

Amount of Rainfall Very Cloudy Partly Cloudy Not Cloudy

Heavy 0 i

Moderate 1 0
Light 2 0
None i 0

In Figure 1.2, the clustering is on a diagonal line going from the upper
right-hand side of the table to the lower left-hand side. We refer to this as
the off diagonal, and when most cases cluster about the off diagonal, we
say that the variables are inversely related. The term negatively related is
sometimes used, but the term inversely related is preferred.

Off diagonal Clustering on a diagonal line that goes from the upper right-hand side
of the table to the lower left-hand side.

Inversely related A condition in which most cases cluster about the off diagonal.
A high score on one variable is associated with a low score on the other.
TEAR ITEITLL CELE ESET NEES IO ERE ETON EES NEO CEE RE NET TS NE RSE A ORNS OMENS EE

If we were initially inclined to believe clouds were associated with a lack


of rain and clear days were associated with rainfall, our hypothesis could
have been worded as follows:

The amount of cloudiness and the amount of rainfall are inversely


related.

Warning: The nature of a relationship, positive versus inverse, is taken


from the logic of the hypothesis and corresponds to the stated diagonals in
a table only if the table is set up so that the main diagonal represents “more
with more” and “less with less.”
How We Reason » 21

Suppose Figure 1.1 had been recast as in Figure 1,3:

Figure 1.3

Amount of Cloudiness
Amount of Rainfall Very Cloudy Fartly Cloudy Very Cloudy
Heavy fl
Moderate
Light
None

Although the categories of rainfall amount remain as before, the other


variable, amount of cloudiness, was entered in the sequence opposite the
way it was done in Figure 1.1. Now the clustering is on the off diagonal, but
the relationship is positive, not inverse. Heavy rainfall is still associated with
very cloudy sky conditions; low rainfall, with clear days. While the general
convention is to set up tables so that clustering in the main diagonal indi-
cates a positive relationship, there are exceptions to every convention. One
must be on guard for such exceptions.
Finally, to reiterate a point made earlier, the positive versus inverse rela-
tionship terminology cannot be used unless both variables have categories
representing amounts of the variables! Suppose we have a hypothesis that
states that an individual’s gender is related to his or her hair color, such that
women are more likely to be blondes than men, and men are more likely to
have dark hair than women. A study of 100 people yields the data shown in
Table 1.10.

Table 1.10

Hair Color

Gender Brown or Black Blonde Total

Male 35) 15 50
Female 15 3D) 50
Total 50 50 100

Since 70 of the 100 people appear to be consistent with the hypothesis


(men with dark hair; women with blonde hair), the hypothesis is verified.
22 STATISTICS FOR THE SOCIAL SCIENCES

Yet, we could not say that hair color and gender are “positively associated.”
There is no quantification in either variable in the sense of the categories
implying more or less of the variable. Male and female are two types of
gender; neither category possesses more or less gender. The same applies
to hair color. Blonde is perceived by most people as a different color than
dark hair, but a blonde has neither more nor less an amount of hair color
than a dark-haired person. (Ignore the fact that physicists do view colors
in amounts—the frequency of light from one end of the spectrum to
the other.) Thus, the term positive or inverse would be unclear in charac-
terizing this relationship and should not be used.
Let us examine how some social science hypotheses might be worded
in this new manner:

»® Income and social alienation are inversely related.


® Occupational status and educational level are positively related.
>» Support for the existing political system and expectations of impend-
ing improvement in one’s living standards are positively related.
» Time spent viewing television and time spent reading print media are
inversely related.
® Income and support for organized labor are inversely related.

We could not word the hypothesis relating religion to partisan identity


in this manner, however. Such a hypothesis relates attributes—Republican,
Democrat—to other attributes—Protestant, Catholic, and so on. These
attributes are not quantitative in nature. They do not represent amounts of
the variables.

ASSOCIATION AND CAUSATION

When we do research, we are ultimately seeking to infer the cause of the


phenomenon we study. What variables, when changing in value, cause the
variable we are studying to change? What factors account for the amount of
rainfall? If levels of cloudiness are associated with levels of rainfall, and we
assume for a moment that the actions of other variables such as tempera-
ture and humidity play no role in the direct relationship between cloudi-
ness and rainfall, we are tempted to suggest a causal relationship between
the two. We might assume that changes in levels of cloudiness bring about
changes in the amount of rainfall. In effect, we are saying that it is the
clouds that cause the rain. Given what is known about meteorology and cli-
mate, it is a logical assumption that clouds cause rain. If we knew nothing
How We Reason & 23

else about weather, however, we might just as likely conclude that it is the
rainfall that causes the cloudiness level. Which of the two directions of
causation we choose will often depend on two things: the logic of the situ-
ation and the temporal sequence of the variables, or which variable came
first in time.
~
TAILLE ESSSS SNELL EEE RIE EEE EELS SE SELLE NNO RLS SILLS SOHN SB

Cause When one phenomenon being studied brings about the other.
NEE SEER

As a point of departure, it should be borne in mind that, in general, it is


illogical to talk about a pattern of causation between two variables unless it
can be first demonstrated that there is association between the variables.
Without association, it is meaningless to consider causation. As demon-
strated in Table 1.4, the same proportion of rainy days exists under cloudy
conditions as under partly cloudy and as under clear conditions. Knowing
sky conditions does not improve our ability to predict or explain rainfall.
Conversely, knowing whether this is a rainy day or not does not help us pre-
dict or explain cloud level. If variation in one variable does not relate to vari-
ation in the other variable, there is no evidence to suggest that either
variable causes, or brings about a change in, the other variable. Thus, with
the exception of a few instances presented in subsequent chapters, associ-
ation is a necessary condition for causation.
Once association has been established, it may be possible to infer a
causal direction between the two variables. A simple tool for doing this,
particularly when we want to explain and not just predict the change in a
variable, is temporal sequence. If one variable changes earlier in time than
does the other variable, we might assume that the former may cause the
latter, In the case of the rainfall problem, we may observe that the buildup
in the cloud level takes place prior to the rainfall. In fact, once the rains stop,
the cloud level often dissipates. Since the cloudiness precedes the rainfall,
we may assume that clouds cause the rain. Since the rain follows the cloudi-
ness in time, it would make no sense to suggest that the rainfall causes the
cloudiness. The cause precedes the effect temporally.
SEE ELEN
MME SENOS
SSTEE ENE MOE ELIE
LUTTE AESSSI ERNEST

Temporal sequence When one phenomenon being studied occurs earlier in time
than the other.
IEEE LEE LEELA DIE ELLIE SHH sere
HL SSSSTORES TEESE CELLS

Unfortunately, we are not always able to ascertain the time ordering of


the variables. The gender versus hair color problem presented in Table 1.10
24 STATISTICS FOR THE SOCIAL SCIENCES

is an example of this dilemma. We receive through inheritance of genetic


factors both our gender and hair color predispositions before we are born.
From this perspective, it would be just as illogical to assume that gender
causes hair color as it would be to assume that hair color causes gender.
Even if there is association found between the two variables, there is no evi-
dence that one either preceded the other in time or could logically have
caused the other.
Given this fact, it would seem reasonable not to concern ourselves with
the causation question at all unless we had evidence for making a causal
inference. Unfortunately, many of the statistical techniques used in data
analysis require that we designate, 7 advance of calculating the statistic,
which variable is doing the causing and which variable is being caused.
Because of this problem, there will be times when we may have to arbitra-
rily select one variable to be doing the causing or explaining and, by default,
assume that the other variable’s changes are being “caused” by changes in
the former. In the case of the gender versus hair color problem, if we are pri-
marily interested in studying hair color, we are likely to assume that for our
purposes, gender “causes” hair color. If our goal were to account for gender
differences, we would have to treat differences in one’s hair color as leading
to differences in one’s gender.
Likewise, in American politics, there has been a relationship between
religion and partisan identity. Although it has been dissipating in recent years,
Protestants have had a slightly higher affinity for the Republicans, whereas
Catholics and Jews have been more likely to vote Democrat. What is the vari-
able doing the causing? Probably religious identity comes first in time,
although not by much, so it would be logical to consider religion the cause
and partisanship the effect. This would certainly make sense to a political sci-
entist who would be trying to account for party identification. However, if
one’s field is the sociology of religion, and religious identity is the subject of
inquiry, it would be perfectly logical to treat partisan identity as causing reli-
gious identity.
When we must differentiate the variable presumed to do the causing
from the variable being caused—whether the selection is based on logic or
based on an arbitrary decision—we usually call the variable being described,
caused, or explained the dependent variable, and the variable doing the
causing or explaining is the independent variable.

——eceseestiesneeenenseenensennesneeteneneennsee
Dependent variable The variable that is being caused or explained.

Independent variable The variable that is doing the causing or explaining.


eR NRE CS HH MAS RE HS RSS OSAP ES SONS
How We Reason ® 25

The dependent variable is the “causee”; the independent variable is the


“causer.” Changes in the dependent variable depend on changes in the inde-
pendent variable but not necessarily the other way around, just as increases
in rainfall depend on increases in cloud level, but rainfall does not cause the
clouds to form.
If we hypothesize that_social inequality leads to revolution, then the
occurrence of a revolution depends on prior social inequality. Revolution is
the dependent variable (being caused), and social inequality is the indepen-
dent variable (doing the causing). If we believe that air pollution (occurring
first in time) causes certain forms of cancer, then level of air pollution is the
independent variable, and the cancer rate is the dependent variable. Implicitly,
we assume in the latter instance that cancer levels do not cause air pollution.
Some people take issue with the use of the terms dependent and inde-
pendent variable in cases where no logical causal ordering can be inferred.
They prefer the term criterion variable as a substitute for dependent
variable and predictor variable as a replacement for independent vari-
able. With this alternate terminology, there is no connotation of causality or
implication of change in one variable being dependent on change in the
other. Nevertheless, while one should be aware of this alternative to the
terms dependent and independent variable, this text will continue to use
the more traditional dependent/independent variable terminology.

seems ee

Criterion variable A substitute for the term dependent variable.

Predictor variable A substitute for the term independent variable.


SOOO SAUL EEE EEE OORT ISIS NEES NSE

THE UNIT OF ANALYSIS

There remains one other item of discussion in this review of the scientific
method—the unit of analysis. The unit of analysis is what we actually
measure or study to test our hypothesis. It is not the variable being studied
but rather the entity being studied—the person, place, or thing from which
a measurement is obtained. In the rainfall problem, days were the units of
analysis. For each of the 30 days, we took two “measurements,” the pres-
ence (or amount) of rainfall and the presence (or amount) of clouds. In the
hair color/gender problem, individual people were the units of analysis.
For each person, we determined two things: that individual’s gender and
that individual’s hair color. In the problem asking whether social inequality
led to revolution, we would have to design a study in which we collected
26 @ STATISTICS FOR THE SOCIAL SCIENCES

information from a number of countries. Thus, country or nation-state


would be the unit of analysis. For each country in our study, we would then
seek to ascertain its people’s level of inequality and also whether revolution
had taken place in that country during some specified time interval. That
would likely be the way a comparative political scientist would handle the
problem. By contrast, a social psychologist might hypothesize that for an
individual, his or her self-perception of being the victim of social inequality
would influence his or her tendency to be supportive of revolution. Here,
individuals, not countries, would be studied, so individuals would be the
units of analysis. An urbanist might assume that where murder rates are
high, so too would robbery rates be high. Data might be collected from a
number of cities, finding each one’s murder and robbery rates for a given
year. Cities would be the units of analysis.

Unit of analysis What we actually measure or study to test our hypothesis: from
whom or from what the measurement is made.

Since so many social and behavioral scientists study aspects of human


attitudes and behavior, we often encounter an individual of one kind or
another being studied as the unit of analysis. Thus, many examples in this
book also use the individual as the unit of analysis. It must not be assumed
that this is always the case, however. Depending on availability of data, the
units of analysis could be individuals; business firms or social organizations;
counties, states, or provinces; or nations or international alliances. Care
must be taken to avoid assuming that what holds true for one unit of analy-
sis holds true for others. Conclusions about the behavior of business firms
do not necessarily carry over to individuals or counties or other units of
analysis.

Example

Suppose we were doing a study of some state’s criminal justice system.


During the process of collecting data, we notice that in the case of public
defenders, those people employed in urban areas seem to have higher
caseloads than those employed in rural areas. Assuming that we wish to
understand caseload (amount of cases per public defender) and that a
community's population size (urban, implying large population, and rural,
implying small) may account for the variations in size of public defender
caseload, then caseload is our dependent variable, and population size is
our independent variable.
How We Reason » 27

The hypothesis we induce from our observations is as follows:

There is a relationship between the size of a public defender’s caseload


and the size of the community where that public defender is employed,
such that public defenders in urban communities have higher caseloads
than their colleagues in rural communities.

Noting that both caseload and population size are quantifiable variables, we
may simplify our hypothesis as follows:

The size of a public defender’s caseload and the size of the commu-
nity employing that individual are positively related.

To test our hypothesis, suppose we have access to data for each county in
that state or province showing the county’s average public defender case-
load and also that county’s population. County is our unit of analysis.
To keep our example very simple, assume that we establish a cutoff
point in terms of caseload and another cutoff point in terms of population
size, such that each county is classified as either high or low in terms of
public defender caseload and urban or rural in terms of population size.
If the hypothesis we induced is assumed to be true universally, it
is assumed to be true—we deduce—for this particular province or state.
We categorize each county in terms of caseload (high versus low) and pop-
ulation (urban versus rural).
Note that our hypothesis is empirical; it can be tested from the data at
hand. Whether or not the hypothesis is true is kept apart—as much as we
can—from our own normative judgments. The facts will hold, regardless of
our Own normative opinions and beliefs about what should be the case.
These normative beliefs could be any of a number of possible attitudes:

Public defenders ought to have equal caseloads, regardless of popula-


tion density, since it is wnfair for urban case workers to have greater
workloads than their country cousins.
Public defenders in cities ought to have higher caseloads than their
rural counterparts since cities have higher concentrations of poor and
disadvantaged, and thus more crime. Therefore, urban personnel
should be working harder.
Public defenders in rural areas ought to have caseloads as high as their
more assertive colleagues in the cities.
Public defenders in cities ought to have higher caseloads than rural
public defenders since the former were foolish enough to select jobs in
the cities.
28 << STATISTICS FOR THE SOCIAL SCIENCES

Regardless of what we think ought to be, the empirical study will tell us what
actually is the case.
Assume that we are studying 40 counties, of which half are classified
as urban and half are rural. A tabulation of our results might look like
this:

Table 1.11

Size of County

Caseload Urban Rural

High 17
Low 3 15
Total 20 20

We note that out of 40 counties, 32 are consistent with the hypothesis


(high load/urban or low load/rural). Only 8 counties are classified in a manner
inconsistent with the hypothesis. The evidence suggests that our hypothesis
is confirmed.
But what about the 8 inconsistent counties? Why did 3 urban counties
have low caseloads and 5 rural counties have high caseloads? To move
toward a theory of caseloads, we need to test the impact of other indepen-
dent variables on caseload. Which ones might they be? Perhaps the 3 urban
counties with low caseloads are prosperous counties with a high tax base.
We could study the impact of median income or tax revenue on caseload.
Perhaps the political party controlling the county’s government has an
impact. We could make the governing party our independent variable.
Perhaps the nature of the crimes committed in the county can explain some
of the inconsistencies. Maybe the 5 rural counties with high caseloads have
a large rate of misdemeanors—easily disposed of by the judicial system.
Perhaps the other 15 rural counties have a larger occurrence of felonies. The
cases may be fewer, but their adjudication might be more complex and time-
consuming. What other independent variables might you want to test to
build our theory?
Over time, if we repeat these studies in other places and keep getting
similar results, confidence will grow as to the validity of our hypothesis. We
have, to the extent that science allows, verified our hypothesis empirically
and moved along the road of theory building.
How We Reason 29

CONCLUSION
In this chapter, we have discussed the scientific method. Scientific rea-
soning is by no means the only way to understand the world. We could
view the world through more traditional ways, such as through theology,
a political ideology, facts or myths generated by our cultural environment,
or simply what those in authority tell us. All of these, however, require
faith in the sources telling us about the world and faith in those who inter-
pret those sources for us. Scientific reasoning also requires faith, but it is
a faith in ourselves and our colleagues. This is a faith based not on outside
authority but on our own ability to collect and interpret data and our abil-
ity to scrutinize the research of others and to be able to reach the same
conclusions they reached.
The scientific method provides us with logical steps for formulating and
testing hypotheses. This thought process parallels the process used in all
scientific research and remains stable. What does not remain so stable are
the techniques of observation and experimentation used to verify hypothe-
ses, research techniques, and the data analysis techniques used in reaching
conclusions.
Research techniques vary with the field of study. In many of the social
sciences, we use some observation and experimentation techniques, but we
also depend quite a bit on survey research through interviews and ques-
tionnaires. Other social sciences such as psychology may use the same tech-
niques but put more emphasis on experimentation.
Although this chapter focuses on the scientific method, the chapters
that follow concentrate on the techniques of quantitative analysis—not so
much on the research design but on the techniques of measuring and
counting for the purposes of analyzing data and showing how the numbers
come to tell us what the facts are. All of these topics are part of the field of
statistics, the study of how we describe and make inferences from data.
Our aim will be to learn how to make the numbers make sense.

SNORE ANNON

Statistics The study of how we describe and make inferences from data.
sremepeomoene

The techniques employed in analyzing data depend in part on the type


of research design used and in part on other factors that we will encounter
in Chapter 2, “Levels of Measurement and Forms of Data.” Before going on,
try working the following exercises.
30 4 STATISTICS FOR THE SOCIAL SCIENCES

EXERCISES :

Exercise 1.1

i Develop a hypothesis appropriate to your major field of interest. Word the


hypothesis using one of the formats presented in this chapter.
Identify the dependent and independent variables. What was the reasoning
behind your decision as to which was which?
Identify logical categories for each variable. Are you measuring amounts or
attributes?
What is your unit of analysis that you will need to study, and how will you
collect the data?
Put together a table similar to the one in the caseload/population example (see
page 28). Simulate the numbers in the table to resemble the results you would
expect to get if your hypothesis is verified.

Exercise 1.2

For the three hypotheses on pages 10-11, repeat Steps 2, 3, and 4 as in Exercise 1.1.

Exercise 1.3

Formulate hypotheses useful for explaining or for measuring progress toward


attaining the following normative goals. For each hypothesis, state the dependent
and independent variables, their categories (or how you will measure amounts of
each variable), and the units of analysis.

L Women and men should receive equal salaries for similar occupations.
2; Housing restrictions on minority communities should be eliminated.
3. Students studying foreign languages should study the languages of the major
linguistic minority groups in their country (e.g., French in Canada or Spanish
in the United States).
Campaign contributions by interest groups or political action committees
should be limited by law.
Chemical and biological weapons should be eliminated.
The death penalty should be applied in a timely manner.
All who are mentally ill should receive treatment.
Use of illegal drugs should be eliminated.
All citizens should receive a college education.
oe
eeAll gun control laws should be repealed.
How We Reason 31

Exercise 1.4
Each of the following hypotheses has a flaw in either its format or its logic. Identify
the flaw and correct the hypothesis.

1. There is a relationship between women and math anxiety such that women
have math anxiety. :
2. Are age and need for social services positively related?
3. There is a relationship between birth weight and smoking such that mothers
who smoke have lower birth weight.
4. There is a relationship between British political parties and support for national
health insurance such that Labour, Conservative, and Liberal Democrats sup-
port national health insurance.
5, Religion and support for church tax exemption are positively related.
6. Liberals run for public office.
7. Communication graduates earn less than public administration graduates and
thus drive cheaper cars.
8. “Right-brained” people are more likely to vote for conservative candidates.

Exercise 1.5

identify the appropriate unit of analysis for each of the following. Be as specific as
possible.

1. Levels of censorship are greater among countries at war than those at peace.
2. Urban areas have higher juvenile delinquency rates than do rural ones.
3. Per capita income is higher in English counties than in the rest of Britain.
4 _ Two thirds of the kindergarten students in Mrs. Smith’s class at Apple Valley
Elementary School were absent last February 7, due to the flu.
1 . Managers are more likely to contribute to charities than are technicians.
6. First-degree murder rates tend to be higher in southwestern U.S. states than in
southeastern ones.
7. Of all NHL teams that year, Detroit averaged the most goals per game.
8. Elvis Presley sold more albums than Jerry Lee Lewis, Bo Diddley, Chuck Berry,
or any other musicians of that era.
9. Tokyo and Mexico City are the largest and second largest world metropolitan
areas, respectively.
10. The United States had more weapons of mass destruction than did Iraq.
WY KEY CONCEPTS ¥
SAAN LR RELI IEE HOSSEIN Ute EY

measurement frequency distributions ratio level of measurement


levels of measurement/ dichotomies absolute zero
scales n-category variables demographic variables
qualitative ordinal level of individual interval
(categorical) data versus measurement data/raw scores
quantitative data individual ordinal data ungrouped frequency
nominal level of grouped ordinal data distribution
measurement Likert scale grouped interval data
attributes versus frequencies class interval
quantities scores closed-ended versus
individual nominal data interval level of open-ended class
grouped nominal data measurement intervals
ETE OTN is LP enna SEAN ENGR ZORA OR EE
CHAPTER

Levels of Measurement
and Forms of Data

V PROLOGUE V¥

I once worked as an assistant dean for a professor of English who was


the dean of our College of Liberal Arts. I admired him greatly, and today,
I honor his memory. But one thing unnerved me. He would occasionally,
in discussing some of the social science faculty members in our college,
refer to them somewhat disparagingly as “the data gatherers.” This tended
to unnerve me since one of my duties was to gather data on enrollments
and students in our college. So I was a “data gatherer” and, even worse, an
untenured “data gatherer.”
So I looked up the word data in the dictionary, and the definition was
information specifically intended to assist decision making or to provide
analysis. That didn’t seem so evil to me, and, in fact, that was what I was
doing when I gathered it. After many a conversation with my boss, I realized
that it wasn’t “data” that bothered him, it was numerical data! He was com-
fortable with qualitative data—Homer, Milton, and Shaw—but he needed
help understanding numbers.
Years later, and many math-phobic students later, I continue to empathize
with those intimidated by numbers and equations. Yet, with a little patience
and a developing self-confidence, intimidated students can overcome their
fears and become competent data interpreters, if not gatherers. Oh, and what
I never told the dean was that as much as he was uncomfortable with numbers,
that was how uncomfortable I was with a lot of literature. I had to read
Shakespeare in my first-year college English course, and it nearly did me in.
I came to appreciate it only years later, seeing it performed on stage. We all
have our phobias.
thle a

Bo
34 4 STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION

To do research using the scientific method, we must first collect data. The
data we collect are actually measurements or counts of characteristics of the
entities we are studying. There are differing kinds of measurements that we
may make in our study. These measurements also fall into categories of
mathematical sophistication known as levels of measurement.
Levels of measurement range from assignments of attributes, such as
identifying a person’s ethnic background, to assignments of numerical
scores. Scores may be naturally occurring, such as a respondent's age, or
may involve some scale developed by the researcher for the measurement
of some characteristic, such as the respondent’s attitude on some social
issue. The level of measurement used by the researcher is important in that
only certain techniques of data analysis should be used on data measured at
specified levels. Thus, the level of measurement of our data determines
what we may or may not do in our analysis of the data.

MEASUREMENT

When we hear the term measurement, we usually think of a very specific


process such as measuring length by means of a ruler or yardstick or mea-
suring weight by means of a scale. We then assign a number corresponding
to that indicated by the measuring device. Thus, an object is determined to
be 10 inches long and 6 inches wide, or someone weighs 61 kilograms. In the
social sciences, measurement indicates the above but also other simpler
actions such as assignment of a person to a particular category of a variable,
such as to Catholic rather than Protestant or Jewish. When we are able to
assign a number with a specific meaning, we are also able to make precise
comparisons. The person weighing 61 kilograms weighs 2 kilograms more
than the person weighing 59 kilograms. In other instances, we may just be
assigning subjects to categories. If one person is an Australian and another
person is a New Zealander, and we categorize them according to nationality,
we are also “measuring” them, but we usually do not use numerical scores.
Even if we coded Australians with the number 1 and New Zealanders with the
number 2, as we often do for computer data entry, the numbers could not be
compared the way weights are compared. Here numbers are stand-ins for
names but do not indicate amounts the way weights do.

Measurement A very specific process, such as measuring length, but also other
simpler actions such as assignment of a person to a particular category of a variable.
meson sats armas ee ehh ERS SHER RENNES EES EES SSS SOMERS LSE RONSON
Levels of Measurement and Forms of Data » 35

Depending on the meaning of the variable chosen and our available


methods of researching that variable, we select a particular pattern of
measurement. The kind of measurement selected will fall into one of four
general categories known as levels of measurement or scales: nominal,
ordinal, interval, or ratio.

Levels of measurement or scales Measurement that falls into one of four general
categories: nominal, ordinal, interval, or ratio.

Qualitative and Quantitative Data

We are using the term measurement to include two somewhat different


things—assignment by category and assignment by amount. When we
assign to categories that do not imply amounts, we often refer to the data as
qualitative or categorical data. This is the case with the first level of mea-
surement, the nominal level, which will be discussed shortly. We assign to
categories based on attributes that do not imply quantities. For instance, if
I asked college students to tell me their major, some would say psychol-
ogy, some sociology, some political science, some anthropology, and so
on. These categories are qualitative categories. They differ by attribute, not
amount. (Psychology is not more of a major than sociology; it is a different
major.)

Qualitative or categorical data Data that are assigned to categories that do not
imply amounts.

At the other end of the spectrum, the interval and ratio levels
of measurement are concerned with amounts. Examples would be age
(measured in number of years) or income (measured in dollars, francs,
rupees, etc.). These are examples of quantitative data. Here the scores do
mean amounts of the variable.
ema

Quantitative data Data that are assigned to categories that are involved with
amounts.

In between the two extremes is the ordinal level of measurement,


where rankings replace exact quantities—Jones has the highest income and
Smith the second highest income. Or, our subjects are placed in categories
that possess a logical ordering—Smith is categorized as high in income,
36 STATISTICS FOR THE SOCIAL SCIENCES

while Jones is placed in the middle-income category. In the latter case, people
are placed in categories, but unlike nominal-level data, the categories have
a logical ranking or ordering. (People in the high-income category earn
more money than those in the middle-income category. )
Throughout this chapter and this book, it will be important to differen-
tiate qualitative from quantitative variables and, indeed, determine the
exact level of measure of any given variable. This is important because the
techniques used to analyze a variable will be determined by that variable’s
level of measurement. As we move up the ladder from nominal to ordinal
to interval to ratio levels of measurement, the techniques of analysis will
be geared to the level of measurement. Something designed to be used
on interval-level data, for example, cannot be used on ordinal-level data;
something designed for ordinal-level data cannot be used on nominal-level
variables. That is why we need the skills introduced in this chapter.

NOMINAL LEVEL OF MEASUREMENT

The term nominal pertains to the act of naming. Here we are assigning our
subjects, our units of analysis, to a particular category of a variable. These
categories represent differing attributes, not quantities.

Nominal Pertains to the act of naming.

Attributes Characteristics that do not necessarily involve amounts, such as one’s


gender, eye color, or religion.

People differ from one another not only in quantifiable ways, such as
weight and height, but also in nonquantifiable ways based on attributes or
characteristics possessed.
An example would be the variable gender, which contains two cate-
gories, male and female. Neither category reflects more or less “gender.”
There are simply two different categories. Jim happens to be male and Jill
happens to be female. Also, there is no particular reason to list the cate-
gories in a particular order. The variable “religious preference” may have
three categories for a particular study: Catholic, Protestant, and Jewish.
There is mo reason to place Catholic first or Protestant second. We could
have as easily said Protestant, Jewish, and Catholic. Each is a different type
of religious preference. The categories reflect differing types of religious
preference, not differing amounts. Also, there is absolutely no reason to
order these categories in terms of what we anticipate will be the ultimate
Levels of Measurement and Forms of Data » 37

number of people in each group. We are not yet concerned with the
question of the size of each category. Thus, each of our variables could be
presented in any order of categorization.

Gender Religion
Female Jewish
Male Protestant
Catholic

We do require that there be a category to accommodate each subject


or person studied and that each subject go into only one category. If, among
the group being studied, there is a Muslim and an atheist, we must either
add those specific categories or add one or more residual categories to
accommodate them. All three of the following sets of categories would meet
our criteria.

A. Religion B. Religion C. Religion


Jewish Jewish Jewish
Protestant Protestant Protestant
Catholic Catholic Catholic
Muslim Other Muslim
Atheist None Other or None

If one of our subjects was a Hindu, Set A above would be inappropriate


unless we added the category “Hindu,” but Set B would be appropriate (we
would put the Hindu in the category labeled “Other”), and so would Set C
(the Hindu would be listed in the category labeled “Other or None”). The
choice of what religions to list would depend on where we are doing the
study. If we were doing the study in New Delhi, our categories might appear
as follows:

Religion
Hindu Sikh
Muslim Christian
Jain Other
Parsi None

In other words, we include the specific denominations that we would logically


expect to find.
38 << STATISTICS FOR THE SOCIAL SCIENCES

Here are some other nominal-level variables that we might encounter:

Region of Birth (U.S.) Race Major


New England African American Anthropology
Middle Atlantic Asiatic Communications
South Caucasian History
Midwest Native American Political Science
Great Plains Polynesian Sociology
Far West Other
None of the Above

An Australian Aborigine now residing in the United States might be placed


in “None of the Above” under Region of Birth and in “Other” under Race.
We often learn nominal categories in a particular order just for the ease
of learning:

Direction Gender Religion


North Men Catholic
South Women Protestant
East Jewish
West

The category ordering helps us learn them but means nothing in terms
of the variable itself. We could list “West” before “North” or “Women” before
“Men” and lose no information.

Forms of Nominal-Level Data

There are two general ways that nominal data may be presented.
In individual form, there is a listing of each individual subject in the study
(often these are persons) and his or her category assignment for one or
more variables. For instance:

Name of Subject Gender Religion Region of Birth (Canada)


Allen Jones Male Protestant Prairie Provinces
Susan Smith Female Catholic Atlantic Provinces
Robert Blondel Male Catholic Quebec
Lisa Goldberg Female Jewish Ontario
‘somes

Individual Data that are presented in a list of each individual subject in the study
and his or her category assignment for one or more variables.
Levels of Measurement and Forms of Data » 39

Although listing these data individually may be a first step for data entry
in a computer or code sheet, rarely are we interested in the information
on an individual-by-individual basis. Generally, the data are presented in
another form known as grouped nominal data. Each category of the vari-
able is listed, and the subjects are not named but are counted (grouped) in
the category into which each subject falls. Tabulation of these numbers is in
the form of a frequency distribution. We list the variable, its categories,
and a frequency column. If for a group of 10 people, 4 are men and 6 are
women, we would display the information as follows:

Gender f=
Male 4
Female 6

i= 10

Grouped nominal data Data that are presented as a category of the variable listed,
and the subjects are not named but are counted (grouped) in the category into which
each subject falls.

Frequency distribution A tabulation that lists the variable, its categories, and a
frequency column.

The letter f stands for frequency, and the letter 7 stands for the total
number of cases. The sum of the fcolumn yields our 7 (4 + 6 = 10). When
you do a study and someone asks, “What is your 7?” that is the same as ask-
ing you, “How many cases are included in your study?”
Two-category variables, such as gender, are often called dichotomies.
Gender is a two-category nominal scale or, more simply, a nominal dichotomy.
When there are more than two categories, we may specify that category
number.

Dichotomies Two-category variables.

Religion (A Four-Category Nominal Scale)


Protestant
Jewish
Catholic
Other/None
40 @ STATISTICS FOR THE SOCIAL SCIENCES

Religion (A Seven-Category Nominal Scale)


Protestant
Catholic
Jewish
Muslim
Hindu
Other
None

Often, if our variable is mot a dichotomy, we simply say it is measured with


an m-category nominal scale. Here, n-category means any number more than
two. We use this simplification since certain characteristics of dichotomies dif-
ferentiate two-category variables from variables with more than two cate-
gories. Thus, the crucial thing we need to know usually is, “Is it or is it not a
dichotomy?” If it is not a dichotomy, it really does not matter to us how many
categories it has. We shall return to the subject of dichotomies a bit later.

n category A term indicating more than two categories.

ORDINAL LEVEL OF MEASUREMENT

The term ordinal refers to order or ordering. This suggests a quantifi-


able ranking from most to least or some other logical sequence or order-
ing of a variable’s categories. When the quantification is by rank order, it
suggests a sequence but not yet an exact amount of a variable. To say
that in terms of population size, China ranks first; India, second; and the
former U.S.S.R., third—this gives us an ordering. But we do not know
how much larger China is than India. We still do not know the exact
population sizes.

Ordinal Involving a rank order or other ordering.

Forms of Ordinal-Level Data

This becomes a bit more complex than was the case with nominal data.
By individual ordinal format, we mean that we rank each individual
subject from highest to lowest along the variable. The variable is such that
not just different attributes but also amounts are implied. Suppose we rank
our subjects from tallest to shortest:
Levels of Measurement and Forms of Data » 41

Name of Subject Height


Allen Jones Tallest (i.e., first tallest)
Lisa Goldberg Second tallest
Robert Blondel Third tallest
Susan Smith Shortest (i.e., fourth tallest)

The rankings do not show us specific heights. We know that Jones


is taller than Goldberg, but we do not know how much taller. Allen could
be 1 centimeter taller than Lisa, or 6 centimeters, or a half meter. Likewise,
we know that Goldberg is taller than Blondel, but not by how much.
When we actually collect our data, we generally collect them in
individual format as above. In fact, we would collect the data at the high-
est level of measurement available to us. So, ideally, we would like to
have the exact height in centimeters or inches for Jones, Goldberg,
Blondel, or Smith. If that is not possible, as would be the case when
we rank only according to our visual observation, we could still order
each one, as above, according to height and assign the rankings of tallest,
second tallest, and so on.
Often, however, when it comes to presenting data in a paper or report
and we have a large number of subjects to be accounted for, we are far
more likely to encounter a format known as grouped ordinal data. Here,
it is not the individuals who are ordered highest to lowest. Rather, individ-
uals are placed into ranked categories, ordered highest to lowest (or lowest
to highest).

ULMER ELLE RSE SSA ER SCM RSE RNR EERE SSS LMT ER LEVEES
MEE SMES SAE REEL LE RS SSUES MIE

Grouped ordinal data Data that present subjects placed into ranked categories,
ordered highest to lowest (or lowest to highest).
SSM UAMUILER SEE EE OEE SSE SMM EESEAMEN

Height f=
Very tall 5)
Tall 7
Medium 10
Short 6
Very short 4
Ws 30

As with grouped nominal data, the f = column (frequency = ) indicates


the number of people in each category of height. The sum of the frequency
42. @ STATISTICS FOR THE SOCIAL SCIENCES

column is 30, indicated by 7, which stands for “number,” meaning the total
number of people. Another example:

Economic Status f=
Wealthy 10
Upper-middle income 20
Lower-middle income 30
Modest income 20
Poor 10
n= 90

The categories represent a sequence from highest economic status


(wealthy) to second highest (upper-middle) to third highest (lower-middle)
and so on down to fifth highest, or lowest (poor).
To be grouped ordinal data, the categories must represent a ranking
(such as high, medium, low) or some other sequence that appears logical
(such as liberal, moderate, conservative), and the categories must be pre-
sented in that sequence. We can completely reverse the sequence, and the
variable is still grouped ordinal.

Economic Status f=
Poor 10
Modest income 20
Lower-middle income 30
Upper-middle income 20
Wealthy 10
i= 90

If, however, the categories are presented out of sequence, which a good
researcher would never do, the variable is no longer ordinal but must be
treated as if it were grouped nominal data.

Economic Status f=
Poor 10
Upper-middle income 20
Lower-middle income 30
Wealthy 10
Modest income 20
i 90

To be grouped ordinal, we would have to reorder the categories back


into their logical sequence.
Levels of Measurement and Forms of Data » 43

LIKERT SCALES
In the social sciences, particularly in survey research involving administration
of a questionnaire, we often encounter an item whereby respondents are
given a statement and then asked their level of agreement.

Statement: UN troops should be permanently stationed in the


Persian Gulf.
Do you: Strongly Agree?
Agree?
Disagree?
Strongly Disagree?

This is known as a Likert scale (named for its inventor) and may be
considered to be ordinal, going from most agreement to least agreement.

Likert scale A scale whose categories are based on the level of agreement with a
particular statement or issue.

Actually, we are tapping both the nature of the respondent’s opinion


(for or against the permanent stationing of UN troops in the Persian Gulf)
and how strongly that respondent feels on the issue. The reasoning
behind treating this as a single ordinal variable is assumption of a logical
sequence. The “strongly agrees” are more in agreement than those who
agree not as strongly. The “agrees” are more in agreement than the
“disagrees,” and “disagrees” do not disagree as much as the “strongly
disagrees.” Therefore, “disagrees” are more in agreement than the
strongly disagreeing respondents. Thus, there is an ordering from most to
least agreement.
What if a respondent is unsure of his or her opinion? The Likert scale is
still ordinal as long as we add an “unsure” response and place it in the middle.

This is ordinal: Strongly Agree


Agree
Unsure
Disagree
Strongly Disagree

The “unsures” are less in agreement than the “agrees” but more in
agreement with the proposition than those who disagree.
44 @ STATISTICS FOR THE SOCIAL SCIENCES

BOX 2.1
Level of Measurement and Dichotomies

Consider the following two dichotomies:


Gender Income
Male High
Female Low

Of the two, gender is nominal, and income is ordinal. We could reverse the
sequence of their categories as follows:

Gender Income
Female Low
Male High

Gender, of course, is still nominal. Income remains ordinal because the


categories are still in their logical sequence. Note, though, that with a dicho-
tomy, there is 70 way to take the categories out of sequence. All we can do is
reverse their ordering, but even in doing that the sequence is not broken.
Consequently, a nominal dichotomy may be treated as if it were an ordinal
dichotomy, even though it really is not ordinal in terms of the logic of its cate-
gories. For this reason, techniques of analysis that apply only to ordinal level
data may also be applied to nominal dichotomies. While we differentiate nom-
inal from ordinal dichotomies, many methodologists do not make that differ-
entiation and subsequently treat all dichotomies as a single category of data.

If, however, we put the unsure response at the bottom (out of logical
sequence), we only have a nominal scale.

Strongly Agree Strongly Disagree


Agree Unsure
Disagree

Those who are unsure are not more in disagreement with the proposi-
tion than those strongly disagreeing with it. So the logical sequence is bro-
ken, and the ordinal nature of this variable disappears.
In studying U.S. politics, we often list the variable party identification in
a way that resembles a Likert scale.

Party Identification
Strong Democrat Strong Republican
Democrat Republican
Independent
Levels of Measurement and Forms of Data j» 45

Although it may be debated, we treat such a variable as being an


ordinal level of measurement by assuming that Democrat and Republican
are logical opposites and there are no other meaningful “third parties.” Since
so few U.S. citizens belong to such “third parties,” we lose only a tiny number
of cases by doing so. Thus, in terms of identification with the Democratic
Party, we assume that the “Strong Democrat” will most often Support that
party’s candidates, a “Democrat” usually will, and an “Independent” may or
may not support the Democrat but is not predisposed either in favor or
against Democrats. The “Republican” is less likely to support a Democrat but
is not as strongly predisposed against Democrats as is the “Strong Republican.”

SCORES VERSUS FREQUENCIES


Up to this point, most numbers that have appeared in this text are
frequencies. They are headcounts or tallies indicating the number of cases
in a particular category or the total number of cases measured. The tables in
Chapter 1 are composed of a series of frequencies. Thus, in Table 1.7, there
were 8 cloudy days on which it rained, there were 17 days having no rainfall,
and so on. In Table 1.10, there were 35 blonde females out of a total of 50
blondes. In this chapter, we found frequencies such as the number of males
versus females or the number of individuals studied whose economic status
could be classfied as wealthy.

Frequencies Headcounts or tallies indicating the number of cases in a particular


category or the total number of cases measured.

When we came to the ordinal level of measurement, however, we encoun-


tered, in addition to frequencies, numbers being used to represent rankings
or scores. Jones ranked first in height, Goldberg ranked second, and Blondel
ranked third. These rankings were not frequencies but rather represented
relative amounts of the variable being measured: tallest, second tallest, third
tallest, and so on. Once we move into the interval level of measurement, we
will encounter whole-number scores, which are more than mere rankings.

Scores Numbers that are used to represent amounts or rankings.

INTERVAL AND RATIO LEVELS OF MEASUREMENT

The next level of measurement is known as an interval level of measure-


ment or an interval scale. Here the subject receives a numerical score rather
than a ranking. Imagine that our subjects undertook an hour of strenuous
physical exercise, after which their body temperatures were recorded:
46 @ STATISTICS FOR THE SOCIAL SCIENCES

Temperature
Temperature (Individual, Interval)
Name of Subject (Individual, Ordinal) 2G: °F

Allen Jones First (highest) 37.55 98.8


Lisa Goldberg Second 37.49 98.7
Robert Blondel Third Seine.) 98.5
Susan Smith Fourth (lowest) oie ai 98.3

Interval level of measurement An interval scale in which the subject receives a


numerical score rather than a ranking and where zero is an arbitrarily chosen point rather
than lack of what is being measured. Scores may be below zero as well as above zero.

The interval data may be added and subtracted. They enable us to


define the distance between subjects’ scores. With rank orderings, we knew
that Jones had a higher temperature than Goldberg but not what the actual
difference was. With interval-level data, we know that since Jones measures
98.8 °F and Goldberg 98.7 °F, Jones is one tenth of a degree warmer than
Goldberg, the person with the next highest temperature. Similarly, Jones is
0.3 °F warmer than Blondel, and Blondel is 0.2 °F warmer than Smith.
Interval levels of measurement correspond to what people generally mean
when they use the term measurement.
Suppose in the problem where we examined the height of our
subjects, we listed each individual’s actual height as well as the rank order.
The height, as measured in centimeters or in feet and inches, would
appear to be measured at the interval level, but it is really an even higher
level of measurement known as the ratio level.

Ratio level A level of measurement similar to interval level, but where zero is an
absolute zero, meaning none of what is being measured. Scores may not be below zero.

Height
Name of Height (ndividual, Feet
Subject (Individual, Ordinal) Ratio) Centimeters and Inches
Allen Jones First 183 6'0"
Lisa Goldberg Second 180 Ske
Robert Blondel — Third 170 eae
Susan Smith Fourth 160 D Oe

Interval and ratio levels are nearly identical. The difference between the
two is the nature of the meaning of zero. In interval data, zero is an arbitrary
Levels of Measurement and Forms of Data p» 47

point, whereas in ratio data, zero is an absolute zero, Meaning a complete


lack of the variable being measured. An example would be temperature.
In the commonly used Fahrenheit and Celsius scales, zero is arbitrarily
selected, and it is possible to have below-zero temperatures. In the Kelvin
scale, however, zero° K (-273° C) is called absolute zero. In theory, it cannot
get any colder than absolute zero, so there are no below-zero readings on the
Kelvin scale. The distinction between interval and ratio levels of measure-
ment leads to other mathematical implications, but for our purposes, these
implications are not important.
In the social sciences, many variables above the ordinal level of measure-
ment are actually ratio level. We encounter few examples where zero is arbi-
trarily assigned. In fact, many statistics books do not even differentiate interval
from ratio levels. We should be aware of the distinction between the two levels
of measurement, but we should also be aware that throughout this text, when
references are made to interval-level data, we really mean both interval and
ratio levels (unless told otherwise), and most examples of techniques requir-
ing interval level of measurement will actually be applied to ratio-level data.
Our levels of measurement may be thought of as increasing in sophisti-
cation, as we move from nominal to ordinal to interval to ratio. At the inter-
val and ratio levels, we assign an exact measurement with a numerical score
indicating the quantity measured. Height (in inches or centimeters), weight
(in pounds or kilograms), income (in dollars or pounds sterling), age (in
years), academic performance (test scores)—all are examples of such data
used in the social sciences.
In the social sciences, we also use interval- and ratio-level back-
ground information on the human subjects studied: age, years of schooling,
income, and so on. These are known as demographic variables. For other
units of analysis, a wide variety of interval scales also exist. Examples include
crime rates for metropolitan areas, counties, or larger political units; other
kinds of census data such as mortality rates; percentage of people in some
region who are employed in agriculture; per capita gross national product;
and so on. In addition, many of us are constantly examining public opinion
and voting statistics pertaining to elections at all levels of government.

Absolute zero A zero that means a complete lack of the variable being measured
rather than some arbitrarily chosen point.
Demographic variables Background information on the human subjects studied.

As well as using existing census or other collected data, many of our


endeavors involve the actual creation of interval- or ratio-level scores. This
is done, as you will see in the next chapter, by creating indices that purport
48 @ STATISTICS FOR THE SOCIAL SCIENCES

to provide interval-level scores for social or political attitudes. These scores


might measure religious tolerance or attitudes toward capital punishment,
abortion, military intervention in the Middle East, censorship, drug use—
whatever issues are salient at that time and place. Such indices are often
constructed from scales of ordinal level of measurement.
We also create interval scales based on examining roll call voting by
elected representatives. What is a particular senator’s roll call voting score
on the issue of free trade? Past studies have examined judicial decisions in a
similar manner. What is a particular judge’s record of decisions on Cases per-
taining to pornography? Voting behavior studies have also been done at the
international level. How often is Israel’s vote in the UN General Assembly
the same as that of the United States? How often is it the same as Egypt's?
Since the interval level of measurement is a more sophisticated level,
we can do more things with interval data than with nominal or ordinal
data. In fact, when we arrive at techniques that manipulate large numbers of
variables at once, virtually all such techniques in use assume interval-level
data. (You will meet some of these in Chapter 14.)

BOX 2.2
Changing Levels of Measurement

Because certain mathematical assumptions are violated when interval-


level data are “created” from lower levels, there is some lingering criti-
cism of this practice. Nevertheless, all the social and behavioral sciences
make use of this.
One should also be careful of going the other way and coding an
interval variable such as income as if it were only ordinal (for example,
high, medium, or low income). This is a good way to present data, but
beyond presentation, statistical techniques should be performed on
the original interval-level variable and not its ordinal clone.

Forms of Interval-Level Data

Interval-level data may appear as individual listings, often called raw


scores, as was also the case with ordinal data. For purposes of computation,
this is an appropriate format since we are best served when calculations are
based on all available scores. For a group of our subjects whose ages are
given, if we wanted to average the scores (later we'll call this average the
Levels of Measurement and Forms of Data » 49

arithmetic mean), we would need to add all scores together and divide
by
the total number of cases:

Name Age
Aaron 6
Bryan ills
Edie 23
Farah 16

Raw score A simple numerical score.

Adding the scores yields a total of 60, and there are 4 subjects (7 = 4).
Dividing 60 by 4, we get 15.0, so the average age for this group is 15 years.
Suppose we had a larger 7, say 30 people. While we (or our computer)
could add up all 30 scores and divide the sum by 30, a listing of 30 scores by
themselves would be difficult to interpret until we had calculated averages
or other measures. By contrast, when we examine the four cases above, it is
not hard to see that we have 4 young people ranging in age from 6 to 23. It
would be harder to ascertain this kind of trend from a larger listing of scores.
For this reason, we often make use of an ungrouped frequency dis-
tribution format for presenting interval-level data to the reader. Here, we
list the scores in sequence (usually highest to lowest), making sure to
include every score that actually appears in our results. (Scores that could
occur but do not actually appear in the final results could be listed with a fre-
quency of zero or deleted from the listing, as we see fit.)

Ungrouped frequency distribution Scores listed in a sequence (usually highest to


lowest) that includes every score that actually appears in our results.

Suppose we have a group of 30 people whose scores along some vari-


Ape are As follows, 29) 26, 26725, 23,22. 21 21821 20820720720, 2099;
HOS O A OedS vie, I tor 16, toed4, 14a. 12 ieand 10; Ghésescores
have been listed in descending order for our convenience.) We can see that
the maximum score is 29, and the minimum score is 10. We begin listing
the scores for our frequency distribution at 29 and list all scores through 10.
(We need not put in scores from 9 to 0 since their frequencies will all be
zero.) Also, although the scores 27, 24, and 13 do not occur in the study
and their frequencies are zero, we list them for purposes of clarity.
50 STATISTICS FOR THE SOCIAL SCIENCES

Score ll

29
28
ay
26
2
24
23
22
al
20
a,
18
17
16
1S
14
13
12
11
10
(eae
ae
ee
ie
ee
Ce
©
eee
ee
n=
1S)=,

Later on, we will consider some techniques for calculating averages


and other statistics from frequency distributions. Primarily, though, the
frequency distribution format is used in presenting results when statistics
have already been calculated for the data.
A third form for interval-level data is known as grouped interval data.
This is an additional simplification of a frequency distribution and is most
useful when, in addition to a large number of cases, there is also a large
number of possible scores, making even an ungrouped frequency distribu-
tion seem unwieldy. The following is an example of how the data in the
above frequency distribution may be grouped:

Score f=
25=29,9 4
20-24.9 10
15=19.9 10
10-14.9 6
Tas 30
Levels of Measurement and Forms of Data ® 51

Grouped interval data | Grouped data that are also at the interval level of
measurement.

Each of the score ranges, 25—29.9, 20-24.9, 15-19.9, 10-14.9, is known


as a Class interval since it indicates the space between two end points.
~

Class interval An interval that indicates the space between two end points.

(Try not to confuse that usage with “interval” level of measurement.)


Each class interval has both a lower and an upper limit (e.g., 25 = lower
limit; 29.9 = upper limit).
The way we group the data can be tricky. Technically, for our grouped
data to remain at the interval level of measurement, two criteria must be
met. First, the class intervals must all be equal in size. To see this, subtract
the lower limit from the upper limit of the topmost class interval: 29.9 — 25 =
4.9. That class interval covers 4.9 units in magnitude. Making the same sub-
traction for the other class intervals shows us that all of them are 4.9 units
in size. Since that is the case, the first of the two criteria for grouped interval-
level data has been met.
By contrast, suppose we had grouped our data this way:

Score f=
25-29.9 a
15-24.9 20
10-14.9 a
= 30

The upper and lower class intervals are each 4.9 units in magnitude, but
the class interval in the middle has a range of9.9 units (24.9 - 15 = 9.9). The
class intervals are not equal in magnitude, so the first criterion for grouped
interval-level data has not been met. What we have now is treated as if it
were grouped ordinal level of measurement, despite the fact that the infor-
mation originated from individual scores and frequency distribution data
that were indeed interval level of measurement. When we grouped the data
into unequal-sized class intervals, we technically dropped down a level of
measurement to ordinal.
The second criterion for grouped interval-level data is that all class inter-
vals must be closed-ended. This means that each class interval must have
both an upper and a lower limit. Consider the following:
52 <4 STATISTICS FOR THE SOCIAL SCIENCES

Score f=
25 and above 4
20-24.9 10
15-19.9 10
10-14.9 6
5-9.9 0
0-4.9 0

n= 40

Closed-ended A class interval that has both an upper and a lower limit.

Notice the topmost class interval: It has a lower limit of 25 but no upper
limit. It is an open-ended rather than a closed-ended class interval. Since
we do not know its upper limit, we cannot assume that it is the same size as
the other class intervals, all of which are 4.9 units in magnitude. Thus, as
presented, this variable is only grouped ordinal-level data.

Open-ended A class interval that has a lower limit but no upper limit or vice versa.

This does not mean that the data should not be presented this way,
particularly if any relevant statistics can be calculated from the data in their
original individual interval-level format. Instead of the people in the top
class interval having scores of 29, 28, 26, and 25, as in the original problem,
so that they are easily grouped into a class interval of 25-29.9, suppose
those four scores were 29, 36, 55, and 72. To preserve equal-sized class inter-
vals, we would have to continue our groupings from 25-29.9 all the way
up to 70-74.9 just to accommodate 4 subjects.

Score f=
70-74.9
65-69.9 0
60-64.9 0
5559.9 1
50-54.9 0
45-49.9 0
40-44.9 0
2 eee ee 1
30-34.9 0
25-29.9 1
20-24.9 10
Levels of Measurement and Forms of Data » 53

15-19.9 10
10-14.9 6
— 0
0-4.9
n= 20

The results are clumsy looking and hard to interpret. In this instance, it is
worth our while to present the data as grouped ordinal, with a top class
interval of 25 and above.
Technically, the lower limit of each class interval must also be provided
for the same reason that we need to know the upper limit. The following—
paralleling the data on the previous page—is also only grouped ordinal-level data.

Score f=
25-29.9
20-24.9 10
15= 10:9) 10
14.9 and below 6
n= 30

We don’t know the size of the lowest class interval.


If, however, in an open-ended lower-level class interval, we can logically
assume that the lower limit is zero, and if making this assumption gives the
lowest class interval a size equal to the other class intervals, we may treat the
data as grouped interval level.

Score f=
25-29.9 4
20-24.9 10
15-19.9 10
oe 1
Below 5 oe:
n= 30

The “below 5” class interval may be assumed to be the same as 0—4.9. From
time to time, we may find class intervals listed as follows:

Score f=
25-30 4
20-25 10
15-20 10
54 ¢ STATISTICS FOR THE SOCIAL SCIENCES

OZ 4
5-10 1
0-5 ak
t= 40

Where would we place a respondent whose score falls at one of the inter-
val’s limits? Suppose someone's score is 25—do we count that score in
the 20 to 25 interval or in the 25 to 30 interval? If the score is exactly 25,
we include that person in the higher (25-30) class interval. If the score
is 25 by rounding off (e.g., 24.9—not quite 25) but we rounded it up to
25, we place that person in the lower (20-25) class interval. This is a
needlessly confusing way of presenting class intervals and should be
avoided!

TABLES CONTAINING NOMINAL


LEVEL OF MEASUREMENT VARIABLES

When one or both of the variables is nominal level of measurement,


care must be taken in interpretation of that table since we may not base our
interpretation on whether there is clustering on a diagonal in the table. This
is because the “ordering” of a nominal variable’s categories is arbitrary.
Suppose we ask respondents whether they agree, are unsure, or disagree
with the statement that a law should be passed severely limiting a woman’s
right to an abortion. The attitude on abortion legislation—our dependent
variable—remains ordinal as long as we keep the “unsure” response in
between “agree” and “disagree.” Suppose, however, our independent vari-
able is religious preference measured by a three-category nominal scale:
Catholic, Protestant, and Jewish. Each of the tables (2.1 and 2.2) is an equally
valid presentation of the data.
Table 2.1 shows a relationship because there appears to be clustering on
the main diagonal, but it is really the change in percentages as one moves
across the row (along the top row, 78.8% to 34.5% to 25.0%) that indicates
the relationship. If we reorder the categories for religion, there is no longer
a discernible clustering on a diagonal, but percentage changes are still
apparent (in the top row, 78.8% to 25.0% to 34.5%; see Table 2.2).
Note that the reordering of religious categories rearranges percentages
for each category of the dependent variable, but in either case, the relation-
ship is equally strong. Since religion is nominal, and therefore the ordering
of its categories is arbitrary, we look for any change in the percentages to
indicate a relationship. We do not examine only the diagonals.
Levels of Measurement and Forms of Data » 55

Finally, note that in Table 2.3, if there are mo percentage changes along
the rows, there is no relationship between the variables. The same percent-
age of each group agrees (30%), is unsure (40%), and disagrees (30%).
Knowing a person’s religious preference, in this case, gives us no additional
aid in predicting or explaining a person’s attitude toward abortion legislation.

Table 2.1 <

Catholic Protestant Jewish

Agree 78.8% 34.5% 25.0%


Unsure li,2 40.5 25.0
Disagree 10.0 PSSA) 50.0

100.0% 100.0% 100.0%

Table 2.2

Catholic Protestant Jewish

Agree 78.8% 25.0% 34.5%


Unsure Wi 25.0 40.5
Disagree 10.0 50.0 ZA

100.0% 100.0% 100.0%

Table 2.3

Catholic Protestant Jewish

Agree 30.0% 30.0% 30.0%


Unsure 40.0 40.0 40.0
Disagree 30.0 30.0 30.0

100.0% 100.0% 100.0%

CONCLUSION

Statistical and analytical techniques are often geared to specific levels


of measurement and forms of data. Thus, to select the appropriate analytical
procedure, we must usually identify each variable’s level of measurement—
nominal, ordinal, interval, or ratio—as well as the form in which that variable
appears—individual (raw score), grouped, or (for interval-level variables)
ungrouped frequency distribution. Some of the exercises that follow will help
you to refine your ability to make these identifications.
56 @ STATISTICS FOR THE SOCIAL SCIENCES

EXERCISES

Exercise 2.1

All of the following variables are grouped. Indicate for each its level of
measurement (see Examples).

Examples: (1) Residence i= (2) Residence f=


House 7 Palace 1
Tent 3 Mansion 3
Apartment ee Hut 25
Total 15 Total 15

Both variables are called “Residence,” although Example (1) is really type of
residence, whereas Example (2) suggests size, status, or cost of residence. Thus,
Example (1) is nominal; differing types of residence are listed, but no sequencing
of the categories is apparent. By contrast, Example (2) is ordinal, going from largest
(and presumably most expensive) to smallest (and least expensive).

1. Cost of Residence i=
Above $1,000,000 3
$250,000-$999,999 5
$100,000-$249,999 i
$ 75,000-$99,999 20
Below $75,000 18
Total 53

2. Cost of Residence f=
$75,000-$99,999 15
$50,000-$74,999 20
$25,000-$49,999 10
0-$24,999 5

Total 50

3. The United States should withdraw all


of its citizens from a known war zone
Response f=
Strongly Agree 25
Agree 20
Unsure 5
Strongly Disagree 5
Disagree 15
Total 70
Levels of Measurement and Forms of Data » 57

. Idealism fe
Very Idealistic 3
Moderately Idealistic 5
Somewhat Idealistic 7
Not Idealistic 4
Total 19

Weight (in pounds) f=


150-200 200
100-150 180
50-100 25
0-50 5
Total 410

Weight (in kilograms) f=


Over 75 200
50-75 180
40-50 1)
0-40 15
Total 410

. Media Censorship ic
Applied to All Topics 25
Applied to Most Topics 25
Applied Only to Military Topics 85
No Censorship 15
Total 150

. Race ‘ f=
Automobile 6
Foot 3
Speedboat 2
Ski 1
Horse 2)
n= 7

. Education j=
Vocational 4
Technical 8
College Preparatory 10
A= 25
58 @ STATISTICS FOR THE SOCIAL SCIENCES

cc
Levels of Measurement and Forms of Data » 59

15. Religiosity (number of times


per year respondent
attends religious services) j=
90-59 0
40-49 5
30-39 20
20-29 50
10-19 20
Below 10 5
Total 100

Exercise 2.2

Look at the following and determine its level of measurement. Low temperature
(°F) on January 1 of last year:

Dayton, OH 10
New York, NY 15
Vancouver, BC AO
Sydney, NSW 70
Fairbanks, AK -10

Despite the fact that this looks somewhat like the examples in Exercise 2.1, it is
really individual interval-level data. The variable is low temperature (the coldest
registered temperature on January 1 of last year). The unit of analysis is city, and the
cities listed are not some variable’s categories. The number to the right of each city
is not a frequency, but rather a score—that city’s low temperature for the day. (See
why it is so important to differentiate frequencies from scores?)
’ For each of the following, indicate the level of measurement and also the
probable unit of analysis.

1. Income
D. Smith $24 000
R. Jones $60,000
M. Jackson $500,000
P. Roberts $15,500

2. Senior Class Ranking


J. Thomas First
P. Roberts Second
A. Albertson Third
C. Chen Fourth
60 << STATISTICS FOR THE SOCIAL SCIENCES

3. Political Party (Canada)


A. Jones New Democrat
S. Smith Conservative
R. Blondel Bloc Québécois
L. Goldberg Liberal
4. The Most Populous Countries (2001)!
Name Population
China 1
India 2
United States 3
Indonesia 4
Brazil 5
Pakistan 6

5. Per Capita GDP—2001' (U.S. Dollars)


Luxembourg $41,950
Norway $37,020
United States $35,200
Bermuda $34,920
Switzerland $34,460
Japan $32,520
Denmark $30,290

6. Women in Parliament'— Percentage of


Seats Held by Women (Lower House)
Sweden 45.3
Denmark 38.0
Finland 37.5
Netherlands 36.7
Norway 36.4

7. Political Ideology
D. Smith Moderate
R. Jones Conservative
M. Jackson Conservative
P. Roberts Liberal

8. Homicide Rate
Chicago Medium
Los Angeles High
Montreal Medium
New York Medium
Toronto Low
Levels of Measurement and Forms of Data » 61

9. Hospital Department Where Employed


A. Jones Oncology
R. Blondel OB/GYN
L. Goldberg X-Ray

10. Conflicts Between Youth Gangs


Arizona Medium
California High
Florida Medium
New York High
Idaho Low

NOTE
1. The Economist Pocket World in Figures, 2004 edition (London: Profile
Books, 2003), pp. 14, 22, and 20.
VY KEY CONCEPTS ¥

demographic data items (on an index content validity


operational or a scale) criterion validity
definitions/working index (scale) construct validity
definitions construction reliability
conceptual validity split-half reliability
definitions face validity test-retest reliability
Defining Variables

VY PROLOGUE ¥
SEEPS LOMIESL EE IELEELELEL SE IIE ES IIE IIE ES IVE EIEN LIDLE LER E SSORES SELESTE ENESE ELISEO LEE SIVAN SHEERS SSI LISLELANL SELES LAREN

Suppose I am a sociologist who wishes to study the level of bigotry in


a designated group of people. Short of asking each one, “Are you a bigot?”
which is likely to be answered in the negative, I would need to come up
with a series of questions, for example, which would tap into the degree
and type of bias—religious, racial, ethnic, and so on—that I might encounter.
I might want to use a system to score the responses in such a way as to
ultimately give each respondent a bigotry score. In addition, I would want
to be sure my questions are actually measuring bigotry rather than some
other phenomenon. The techniques presented below will assist me in
designing my study.
‘EAU
LLU EERO SILI
ENE EEN EEE RR IIE IEEE INNIS IEE DERE SLE LMUES BEERLCDIELIE
RL ESTELELELEL LIER EDEL TELE EE ELLE
MEE LEE
64 <@ STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
In this chapter, we are dealing with the way in which we develop systems of
measurement for the variables we are studying. We begin by determining
how we will make a measurement and what specific criteria we will use for
assigning our subjects or respondents to specific categories of each variable.
Attention is given to the creation of numerical scales or indices of opinions
or attitudes and how we determine the validity and reliability of such scales.
Selected examples of variable measurement are also presented.

GATHERING THE DATA

Let us assume that a researcher has identified one or more hypotheses to be


tested in a study. The selection of the hypothesis also generally locks a
researcher into studying a particular unit of analysis. If the generalizations are
about individual behavior or attitudes, we normally choose human subjects
as our units of analysis. If the hypothesis refers to characteristics of cities,
cities or metropolitan areas become the units of analysis, and so on. In each
hypothesis, there are usually two variables to be studied, a dependent vari-
able whose variation the researcher is trying to explain or predict and an
independent variable that hypothetically accounts for the change in the
dependent variable. To study these variables, we must be able to measure the
characteristic or amount possessed by each of our units of analysis. Before
we can make measurements, however, we must determine exactly what we
want to measure and how we are going to take these measurements.
Suppose we have chosen individuals as our units of analysis and we
intend to administer a questionnaire to each of our subjects. Also suppose,
as is often the case, that we intend to begin our survey by asking each
respondent for some basic background information. We would not want the
subject’s name if we wanted an anonymous survey, but we might like to know
such things as the subject’s age, sex, marital status, religion, ethnic back-
ground, and so on. These social characteristics may be related to many of the
variables in our study. Such background information is called demographic
data. If one of our hypotheses is that liberalism and age are inversely related,
one of our variables—age—will be among the demographic data in the early
part of our questionnaire. For our purposes, what do we mean by age, and
how do we propose to measure it? Do we want the respondent’s age at
the time he or she fills out the questionnaire? Suppose the survey is to be
conducted over several weeks and to several groups of people. We might
select to have the respondent indicate his or her age as of some specific
date; for example, “How old were you as of February Ist of this year?” Let us
Defining Variables p» 65

assume that we are satisfied with the respondent’s age at the time the survey
document is filled out. Are we satisfied to know the respondent’s age only in
years? This is usually the case, but there are instances when we might opt for
more specific information. In studying children of elementary school age, for
instance, we might want the age in years plus months if we have reason to
believe that, for example, a 7-year-old child may respond to certain items
quite differently from a 72-year-old child.

BLOT
U RBIS LOBES
LE DEAREST OCTET romaine

Demographic data Background information that gives the social characteristics


of a subject.
ARE LEED INESSSRN EIDONOE BE EON EI TS EIEN OT A NAS

Suppose for our study that age in years only is sufficient. How shall
we get our age data? With an adult respondent, we may simply use the
following format:

Age: years old

The respondent just fills in the blank with the appropriate number of years.
Most of the time, this is adequate for social research, but imagine a
situation in which we have reason to suspect that the respondent may
misrepresent his or her age. We might want to obtain the age from docu-
mentation provided by the respondent, such as a birth certificate. What if
someone said he was 18 years old, but his birth certificate indicates that he
is only 17% years old? The age we record depends on what we have decided
in advance. If we had decided to accept whatever age the respondent gave,
then this person will be listed as 18 years old. If we wanted the age as indi-
cated on the birth certificate, we would record 17”.
Likewise, in studying voting behavior, we often find instances of people
claiming they had voted in a particular election when they had not. (After
all, we learn that voting is a civic duty.) In this case, the respondent’s answer
to the question of having voted in that election may be a far less accurate
operational definition than one requiring the researcher to confirm the
answer by examining public voting records.

OPERATIONAL DEFINITIONS

In making such decisions, we are formulating a working definition or


an operational definition. For demographic data from an adult sample,
we are usually satisfied to operationalize these concepts by accepting
whatever response the subject provides. We would list the subject as
66 @ STATISTICS FOR THE SOCIAL SCIENCES

18 years old because that was what the subject said, and we assume that he
or she is telling the truth. The operational definition is thus a measurement
definition. It defines how we are going to measure someone or something
to determine the subject’s score on a variable.

Working or operational definition A definition of the way that someone or something


will be measured to determine the subject’s score on a variable.

The idea behind the operational definition is that once formulated and
applied, there would be no disagreement as to the respondent’s score or
category assignment. In a particular room, some occupants might find the
temperature too hot, whereas others are comfortable. Because there is dis-
agreement among the occupants, we cannot characterize the room tempera-
ture as being either too hot or comfortable. Suppose, though, that we agree
in advance to measure room temperature with a thermometer and opera-
tionally define “too hot” to be any temperature equal to or greater than 78° F.
If the thermometer reads 77° F, we consider the room to be comfortable even
though several occupants feel it to be too hot; if the thermometer reads 78° F,
we consider the room to be too hot even though several occupants consider
the room to be comfortable. Thus, the operational definition, by virtue of
its arbitrary specificity, eliminates for our purposes any disagreement as to
whether or not the room is too hot. The disagreement comes in advance of
our measurement when we decide arbitrarily that 78° F is our cutoff point.
When we move from demographic concepts to other social or political
variables, the problems of operationalization may become more difficult. All
of these must be addressed before we can continue our study.
In the case of research involving human subjects, we are likely to face
conflicts between attributes (what we say we are), attitudes (the way we
actually feel), and behaviors (what we actually do). Suppose ideology
(liberal to conservative) is our variable. We could ask the respondent for a
self-assignment to an ideological attribute as follows:

Do you consider yourself to be: (please indicate)

a liberal?
a moderate?
a conservative?

Suppose the respondent checks liberal. We then ask a series of ques-


tions designed to tap attitudes that would reflect ideology, such as attitudes
Defining Variables » 67

on abortion, aid to antidictatorial insurgencies in Latin America, censorship


of “adult” magazines, and so on. Suppose the same respondent who said he
or she was a liberal then gives consistently conservative responses to these
attitude questions. Assuming that we have used attitude questions that
reflect current major differences between liberals and conservatives so that
our questions are valid, we obviously have a conflict between the respon-
dent’s self-assigned attribute (liberal) and that person’s political attitudes
(conservative). In designing our study, we need to know what will be most
germane and useful to us—the attribute or the attitude.
A similar conflict between an attribute and a behavior could occur. Take
a respondent who, when asked his or her political party identification, says
Democrat. We then discover that in the last five elections, the same respon-
dent consistently voted for the Republican candidate. What will be most use-
ful for us in our study, to assign that subject by attribute (Democrat) or by
behavior (Republican)? Subjects are rarely as consistent as we would like them
to be, particularly when they do not perceive the topic that interests the
researcher as having much direct importance in their own daily lives. Because
we must live with such inconsistencies, we as researchers must make deci-
sions about what we are trying to find out and what we will do with the infor-
mation. If, for instance, our goal is to predict a respondent’s vote in the next
election, that person’s past voting behavior is likely to be a better predictor of
future voting than is the self-assigned attribute of party identification.
A related problem in forming our operational definition is that before
we operationalize, we must have consensus on at least the major parts of
our conceptual definition. The conceptual definition is the more general
definition of that concept such as one would find in a textbook or dictionary.
As an extreme example, note the term democracy. As we currently use it in
the West, a democracy is a political system that governs based on a popular
consent determined by free elections. By our standards, prior to reunifi-
cation, West Germany was more democratic than East Germany. Yet, the
formal name for East Germany was the German Democratic Republic, and
at least to a Marxist ideologue, East Germany was democratic in that the
representatives of the workers and peasants, through the Communist Party,
controlled the government. Clearly, we have two very different views and
definitions of the word democracy. Any operational definitions that stem
from the Western concept of a democracy will be far different from the
operational definitions based on Communist interpretations.

Conceptual definition A general definition of a concept such as one would find


in a textbook or dictionary.
68 << STATISTICS FOR THE SOCIAL SCIENCES

The above case is extreme. More commonly, there are agreements as to


the general, conceptual definition but disagreements as to what aspects of
that conceptual definition compose the essence of the concept essential to
the operational definition. An example is the attempt to operationally define
a concept such as freedom. Is a particular country free (its citizens possess
freedom) and, if so, how free? Suppose we begin by looking up the dictio-
nary definitions of freedom and selecting the portion of those definitions
most germane to political freedom.

Freedom: Possession of civil rights; immunity from arbitrary exercise


of authority,’

There are two general parts to the definition: (1) civil rights and (2) exercise
of authority. Should our operational definition be based on one of these?
Which one? Or should we use both?
Suppose we decide to include possession of civil rights. What is a civil
right, and which rights should we include in the operational definition? Civil
rights are rights granted to an individual based on citizenship or national
residency. We might begin with the “four freedoms” in the First Amendment
to the U.S. Constitution:

» Freedom of religion
> Freedom of speech
>» Freedom of the press
> Freedom of assembly

To this list we could add other civil rights gleaned from the U.S.
Constitution’s Bill of Rights:

> The right to bear arms


» Freedom from unreasonable searches and seizures
P The right to a jury trial
» Freedom from double jeopardy
>» Freedom from cruel and unusual punishment

If we examine other documents such as the UN Charter or other bills of


rights, we could add additional items such as the rights of certain linguistic
groups to have their language used as an official national language or the
rights of citizens to a specified economic standard of living.
What we include in our operational definition will reflect our individual
values and levels of knowledge. Once we agree on what to include, further
clarification must be undertaken to tighten our definitions. Suppose we had
decided to use the four freedoms of religion, speech, press, and assembly.
Defining Variables » 69

We still have to clarify what these mean. In the U.S. Bill of Rights, for instance,
freedom of religion really referred to the government’s not making laws
establishing a particular religion. In modern times, many nations have
“established” religions, even though they are, by our definition, democra-
cies (examine the status of the Church of England in the United Kingdom).
The real issue for us to examine is not whether there are official religions in
a country but whether adherents to the other religions are restricted in their
freedom of worship or in other civil rights.
A second consideration is that all freedoms are limited even in the most
democratic of countries. For example, your religion may believe in ritual
human sacrifice, but that does not mean that the state allows you to practice
that ritual. Likewise, freedom of speech is limited. Recall Justice Oliver
Wendell Holmes’s dictum that freedom of speech does not give one the
right to shout “Fire!” in a crowded theater. We limit freedom of the press
through libel laws and anti-pornography legislation. We limit freedom of
assembly by requiring permits to hold public meetings. Therefore, our oper-
ational definition cannot be so tight as to disallow these kinds of limitations.
A final but crucial problem in forming operational definitions is whether
there exist available data that will enable us to code each country in terms of
the specific civil liberties chosen for inclusion in our operational definition. Is
there any source of data available to us that would enable us to determine, say,
the existence and level of freedom of assembly in each country? Economic
and social statistics are available from several sources, but do they contain the
information we need? In the case of our civil rights scores, we may have to rely
on the opinions of experts who are asked to score each country for which
they possess expertise in terms of the freedoms we have included. Some
examples of operationalizing such variables will be discussed later.

INDEX AND SCALE CONSTRUCTION

For attitudinal variables, the operational definition usually is based on a


subject’s response to one or more questions designed to tap the variable
being studied. In a previous example, we determined one’s attitude toward
abortion using a Likert-type response set.

Statement: Abortion should be illegal.


Response: Strongly Agree
Agree
Unsure
Disagree
Strongly Disagree
70 < STATISTICS FOR THE SOCIAL SCIENCES

We could code each response as an ordinal ranking from (1) strongly agree
to (2) agree and so on to (5) strongly disagree, thus creating a rank order-
ing on opposition to abortion. By simply reversing the rankings, (5) strongly
agree to (1) strongly disagree, we would have a rank ordering on support
for abortion rather than opposition to abortion as originally ranked.
A variation on this idea is a (adder question.

Image a ladder on which those |__| 1 Most opposed


most Opposed to abortion stand on |
exesl) 42
top run and those least opposed 3 Unsure
stand on the bottom rung. Where ie 4
on the ladder would you place | | 5 Least opposed
yourself?

The respondent self-selects his or her place on the ladder, and the researcher
codes that response by indicating the number (rank) of the rung chosen.
A second variation is a feeling thermometer. Instead of a ladder, the
subject sees a picture of a thermometer ranging, for instance, from 0° to 100°.
The accompanying statement asks the respondent to self-assign his or her own
“temperature,” with 100° most opposed, 50° unsure, and 0° least opposed.
Such questions may suffice to measure attitudes along single issues. A
problem arises when what we are measuring is a compound variable made
up of many differing attitudes. Suppose we want to measure an individual’s
social conservatism. While in its broadest sense, conservatism relates to
mistrust of change, in the social context, we associate conservatives as tak-
ing certain positions on issues. Instead of asking the respondent to simply
indicate whether he or she is conservative, we might better tap the issue by
asking a series of questions, each designed to tap a separate aspect or
dimension of conservatism. The issues chosen must be carefully selected to
be meaningful in the current social and political context because attitudes
change over time. Forty years ago, many, if not most, conservatives opposed
mandatory desegregation of racially separate schools in the U.S. South.
Today, few conservatives would be opposed.
Suppose we decided on five items (questions) that we considered
good differentiators of conservatives from liberals in contemporary U.S. pol-
itics. The respondent would provide a Likert-type (strongly agree through
strongly disagree) response to each of the following items.

1. Abortion should be illegal.

2. “Family Values” should be taught in schools.

3. Full funding for the Defense Department is needed for national security.
Defining Variables » 71

4. Educational and welfare issues should be primarily handled by the


states or localities, not by the federal government.

5. The controlling or outlawing of handguns by the government is wrong.

As worded above, we would expect a very conservative individual to respond


“strongly agree” to most of the five items, whereas a strongly liberal, least con-
servative individual would “strongly disagree” with most of the items. Thus,
we could assign points for each item’s response and then sum the points for
each item, giving us an index (also called a scale) of social conservatism.

Items The various components (e.g., abortion, family values, etc.) used to
generate a scale or index.

Index A range of scores, treated as interval or ratio level of measurement,


measuring some phenomenon. In this example, the higher one’s score, the more
politically conservative he or she is.

Suppose we score each question’s response as follows:

Strongly Agree 20 points


Agree 15 points
Unsure 10 points
Disagree 5 points
Strongly Disagree O points

If a respondent gave the most conservative response, strongly agree, to


each of the five questions, that respondent would score 100 (20 x 5 = 100)
on the index. The person strongly disagreeing with each item would receive
a total score of zero (0 X 5 = 0) on the index. One who is completely unsure,
assumed to be in the middle on all five items, would score 50 (10 x 5 = 50).
These scores on our social conservatism index (or social conservatism
scale) are treated, for purposes of data manipulation, as an interval level of
measurement.
We would do several other things to refine our index before using it for
actual research purposes. First, we initially set up our questions so that the
most conservative response to each item was strongly agree, but in doing so
we may have introduced bias into the response set. After several questions,
the conservative respondent might automatically answer strongly agree
or agree without carefully reading the question. To avoid this, we “reverse”
some of the questions so that at times the most conservative response
72 4 STATISTICS FOR THE SOCIAL SCIENCES

would be strongly disagree instead of strongly agree. In such instances, the


most conservative response will still receive 20 points, even though it was
strongly disagree rather than strongly agree. In the following example, we
reword two of the items, show the possible responses, and indicate (in
parentheses) the number of points we will assign. In the actual question-
naire, the number of points for each response should not appear in print,
but the ones coding the scores later on would use the point values to
determine the final index score for each subject.

Directions: Circle the response to each of the following questions that most
closely reflects your own opinion.

1. Abortions should continue to be legal.

Strongly Agree Agree Unsure Disagree Strongly Disagree


(0) 6) (10) (15) (20)
2. “Family Values” should be taught in schools.

Strongly Agree Agree Unsure Disagree Strongly Disagree


(20) (15) (10) (5) (0)

3. Full funding for the Defense Department is needed for national security.
Strongly Agree Agree Unsure Disagree Strongly Disagree
(20) (15) (10) (5) (0)
4. Educational and welfare issues should be primarily handled by the
federal government, not the states.
Strongly Agree —_Agree Unsure Disagree Strongly Disagree
(0) (5) (10) (15) (20)

5. The controlling or outlawing of handguns by the government is


wrong.
Strongly Agree Agree Unsure Disagree Strongly Disagree
(20) (15) (10) 6) (0)
The very conservative respondent will answer strongly agree to items 2, 3,
and 5 (for a total of 60 points) and answer strongly disagree to items 1 and
4 (for an additional 40 points). The grand total will still be 100 points for that
individual.
Defining Variables » 73

BOX 3.1
Interval-Level Scores From Ordinal-Level Data

These scores on our social conservatism index (or social conservatism


scale) are treated as an interval level of measurement for purposes
of data manipulation, even though the Likert response set for each
item is really only ordinal. We arbitrarily assigned the point spread and
arbitrarily assumed that the difference between each adjacent response
would be worth 5 points (strongly agree: 20 points — agree: 15 points
equals a differential of 5 points). We have no evidence to verify that
these points reflect the true amount of difference between the two
responses. We did violate some mathematical assumptions in creating
an interval level of measurement index out of ordinal components, but
as previously indicated, this is common practice in the social and behav-
ioral sciences. While our index was developed from only five questions,
most such indices contain many more items than five. The more items
we add, the more possible options of opinion we add to our index, and
the closer our index gets to being truly interval-level data.

VALIDITY

Once our questionnaire is reordered, we would pretest it on a group of


subjects, administering it once and possibly readministering it to the same
group several weeks later. During this pretesting phase, we would be seeking
to refine the scale by determining two things—the validity and reliability of
our questionnaire as a measurement device.
Validity is the extent to which the concept one wishes to measure
is actually being measured by a particular scale or index. Does the scale
measure the concept it claims to measure? Is it congruent to the generally
accepted definitions of the concept? For instance, if occupational income
alone is being used as a measure of poverty, those with low incomes will be
considered to be poor. In most instances, the measure is valid, but what
about the millionaire who does not need to work and therefore has no
income? This individual is not poor by anyone’s definition. Thus, work-related
income is not necessarily a valid index of poverty.

Validity The extent to which the concept one wishes to measure is actually being
measured by a particular scale or index.
74 STATISTICS FOR THE SOCIAL SCIENCES

There are several strategies for determining a measure’s validity.


The first two—face validity and content validity—rely on the internal logic
of the measure. Face validity is the extent to which the measure is
subjectively viewed by knowledgeable individuals as covering the concept.
For instance, my conservatism scale developed earlier in this chapter
seems valid to me. Each of the five items seems to tap a relevant distinc-
tion between more and less conservative people. If Ishowed the scale to
others with knowledge of the subject matter and they confirmed that each
item measured conservatism, I could say that the measure had face valid-
ity. If there was controversy about some item, say, the abortion question,
I would have to ask if in reality the abortion stand was a valid aspect of
conservatism.

LASSE SURED SEES SIOSSSA CONIA EENES TIEN EEENLIST


OY EIEN TET OEE T TET IS

Face validity The extent to which the measure is subjectively viewed by


knowledgeable individuals as covering the concept.
nce

Content validity is related to face validity, being based on logic and


expertise. It asks whether the measure covers all the generally accepted
meanings of the concept. What if Ishowed the conservatism index to my
judges, and they responded that each item had face validity but that the
scale was incomplete? Several of my experts say, “What about communism?
How can you measure conservatism without asking the respondent about
communism?” If we concur that this item must be included in the scale to
give it content validity, then I would need to add a statement such as this:
“Worldwide aid for anticommunist insurgents should be increased.”

LOSE EMEA EERE ALLAN


LON S snmoninanaus

Content validity The extent to which the measure covers all the generally accepted
meanings of the concept.
ssonrepnn oneness
saomesiiite artinggi pSSomNNSA
oo NIN

Two other types of validity are less subjective and more empirical. They
are known as criterion validity and construct validity.
Criterion validity is based on our measure’s ability to predict some
criterion external to it. The criterion could be in the present and currently
predictable (concurrent validity), or it could be in the future (predictive
validity). For instance, suppose we have designed a scale for determining
whether an individual would be good in a management position with a
firm. We can look at those who later became managers and compare their
performance evaluations with their scale scores. If the index has criterion
Defining Variables » 75

validity, those scoring high on it would also be expected to perform well as


managers. If some aptitude test claims to measure mathematical aptitude,
we would expect those receiving high scores to also earn higher grades in
math. If the opposite situation should result, high scores and low grades, or
if those with both high and low aptitude scores performed equally well in
class, then the aptitude test would be a poor predictor of performance and
would lack criterion validity.

SESE RU ESSE EEE SESSILIS EILEEN EEREES WCC EEE EEE EHOMCRSECT LOLCat MiSitsteieoNit eon iC RN

Criterion validity The extent to which the measure is able to predict some
criterion external to it.
POLE EES SSSI BEEEE LEER LL IRE EEEES SOA ELAMEEEEEE SS LEEELEEE ENATERED EEE ELEVEN ELEM NEN AAEM NELLA

Construct validity has to do with the ability of the scale to measure


variables that are theoretically related to the variable that the scale purports
to measure.

SALA EES cE BDU EER SESE EET NEES LEU EEL EEEIEE DORE EE

Construct validity The ability of the scale to measure variables that are
theoretically related to the variable that the scale purports to measure.
eA SONIA LLLLL SSSEESESSE SCOOT OTLOE LE RENEESSELTE SR NEE ESENELSON EEE SNELL ELLE AEE LDL LEELA ELEDVDLALLELLELDLLEELEELLALE AOE ALLEL AED,

Imagine that you have developed a scale to measure overall life


Savistiction, Ihe higher the score, on. the index, the greater is. the
person’s life satisfaction. To establish construct validity for the scale, ask
what characteristics are likely to be related to overall life satisfaction. For
instance, a satisfied individual would be less likely to be a heavy drinker
or a spouse or child abuser. Is this the case with those scoring high on
your life satisfaction scale? If these or other theoretical attributes are asso-
ciated with life satisfaction, we should be able to empirically test the rela-
tionship between one’s score on your scale and alcohol consumption or
incidents of abuse. If these associations are indeed found to be the case,
then your measure is likely to be a valid index of life satisfaction. You have
established construct validity.
Both the criterion validity and the construct validity may be measured
using techniques similar to the association and correlation measures pre-
sented in later chapters of this text.

RELIABILITY

For a measure to be reliable, it must be free of measurement errors.


That is, (a) the overall score should correspond to the scores of its
76 @ STATISTICS FOR THE SOCIAL SCIENCES

components, a type ofinternal consistency, and (b) if the measure is taken


over intervals of time, the scores of individuals should remain consistent
over time as well. Think of an observed score as differing from the true
score due to errors in measurement. That is, the observed score equals
the true score plus or minus some measurement error. Ideally, true relia-
bility is attained when measurement error is eliminated. More realistically,
a score is reliable when we have minimized the impact of measurement
error as much as possible.

Reliability The likelihood that the scale is actually measuring what it is supposed
to measure.

Split-half reliability is one way to measure internal consistency. To


see if the items are all measuring the same concept, we split our overall scale
into two scales, each containing half the original items. Suppose our origi-
nal index was a 20-item scale designed to predict whether a teenager was
prone to juvenile delinquency. We break the 20-item scale into two 10-item
scales either by putting the odd-numbered items in one group and the even-
numbered items in the other or by assigning 10 of the items at random to
one group and putting the remaining 10 items in the second. Then we com-
pare scores by subscale. A person appearing prone to delinquency on one
subscale should also appear prone to delinquency on the other. If this is the
case, we may assume that the original 20-item scale is reliable in terms of
internal consistency.

Split-half reliability A measure of internal consistency that splits an overall scale


into two scales, each containing half the original items.

The second kind of reliability is test-retest reliability, also known as


reliability over time. Reliability in this context has to do with an individual’s
consistency in responding the same way to a specific item over time.
Suppose we were to administer the conservatism questionnaire twice to the
same group of people and compare each set of responses. If the responses
remain about the same over time, that scale is considered reliable. If
responses change, the scale may not be reliable. The cause for unreliability
may lie in the fact that one or more questions were vague or confusingly
worded. As a result, the reader’s interpretation at the second reading may
have differed from the initial interpretation of the same item. For example,
Defining Variables » 77

suppose the question on the conservatism measure about regulating hand-


guns showed that many people opposed handgun legislation the first time
they filled out the questionnaire but showed changes in their response to
support legislation the second time they filled it out. Or the responses may
have changed from favoring to supporting the legislation. Under normal cir-
cumstances, we would consider the item unreliable and delete it from the
final questionnaire, concluding that gun control attitude is not a consistent
and reliable indicator of political conservatism.

Test-retest reliability A measure that determines an individual’s consistency in


responding the same way to a specific item over time.
ERNE
CR RIESE ETS

Before concluding unreliability, be sure that no intervening event


occurred between the first and second administration of the questionnaire
that would cause a consistent one-directional shift of opinion. For instance,
every time there is an attempted or successful assassination of a popular
public figure, attitudes favoring gun control legislation increase. In such a
case, the consistency of the response changes suggests that the item may
still be a reliable indicator of social conservatism.
To be a good measure, a scale or index must be both valid and reliable.
This is often a difficult order to fill given the fact that many social science
concepts are difficult to define and, once defined, are subject to measure-
ment and other human error. In addition, as Babbie (1989)* has pointed
out, there is a certain tension between validity and reliability. Validity seeks
to be inclusive, extending a measure to cover all of the meanings and
nuances of the concept in question. Reliability tends to exclude nuances and
multiple aspects of a variable so as to focus on what can be specifically
scored. One solution is to create several measures for the same concept and
see if they produce similar results.

CONCLUSION

Understanding the nature of operational definitions and formulating actual


operational definitions are among the hardest tasks for students to master.
While in many courses, our task is to broaden the scope of a definition to
include more and more nuances and examples, here our task is to narrow
that definition, making it ever more specific. In many ways, this task paral-
lels the formulation of specific legal definitions. When members of a jury
determine facts and thus the guilt or innocence of the accused, they base
78 < STATISTICS FOR THE SOCIAL SCIENCES

their determination in part on the judge’s instructions, which include the


legal definitions of the charged crimes. For example: What constitutes mur-
der? How do first-degree murder, second-degree murder, and manslaughter
differ under current state law? What facts must be proven for a jury to con-
clude a guilty verdict? The legal definitions given to the jury are as specific
as possible, and the jury members then determine if the facts presented to
them match those required by the definitions.
An expert coder, like a juror, can take the researcher’s operational
definition and conclude on a case-by-case basis whether the definition is
met. For example, based on the operational definition supplied, the coder
can determine whether or not country x is economically developed. But
suppose we have no coder. How do we then decide whether country x is
an economically developed country? One way is to make use of one or
more variables as stand-ins for, or indicators of, economic development. We
would pick specific variables, each of which may tap only part of the con-
cept of economic development, such as percentage of the population in
agriculture, radios per 1,000 population, and so on. Finding such data and
selecting valid and reliable indicators of the concept we want to measure
are often not easy.
Furthermore, an improper operational definition will lead to improper
statistical results because the statistics will be only as good as the data. In
popular terminology, this is the GIGO principle: “Garbage in; garbage out.”
Following are four general situations that lead to misleading operational
definitions.

1, Ideological Assumptions. An aspect of the operational definition may be


a debatable ideological assumption. For instance, there is an organization
that rates each country on its adherence to principles of human rights, par-
ticularly its treatment of prisoners. In its rating system, capital punishment
is considered to be an indicator of reduced human rights. Several countries,
including the United States, get reduced ratings because they have capital
punishment.

2. Situational Factors. Situations specific to the subject lead to misleading


conclusions. For example, a researcher studying levels of freedom in various
countries gives a certain country a low score because it is practicing censor-
ship. The researcher does not take into account the fact that the country
was at war at the time of the study. Thus, censorship of militarily sensitive
subjects had been instituted, whereas in peacetime there would have been
Defining Variables j» 79

‘no censorship. Another example would be a recent immigrant with a low 1Q


score. The reason for the score being low was that the IQ test administered
was not in that person’s native language. Thus, it was the testing situation,
not the person’s intelligence, that led to the low score.

3. Key Word Inconsistency. Respondents identify incorrectly with


popular terms. For example; students were asked to assign themselves to
one of three categories: liberal, moderate, or conservative. Later, they
responded to items dealing with policy issues normally thought to differ-
entiate liberals from conservatives. Several who had identified themselves
as Conservatives responded to the specific policy items with clearly liberal
preferences.

4. Poor Predictability, The operational definition has a poor track record


in predicting what it claims to predict. For instance, many high school
students in the United States take standardized aptitude tests to deter-
mine their probable performance in college. These test scores are often
used as criteria for college admission and for qualification for varsity sports.
Yet actual studies of the relationship between aptitude test scores and
first-year college grade point averages (GPAs) show that only about 6% of
the variation in grade point average can be accounted for by such aptitude
test scores. (It has been argued, though, that this is because not all who
take the tests actually attend college—only those who score high enough.
That may have been the case in the past, but today almost everyone
can get admitted to some 2- or 4-year college in the United States. It would
be interesting to see if the predictability of GPAs from such scores is
now going up.)

In all of the above examples, weaknesses in the operational defini-


tions could lead to misleading statistical results. One must guard against
such pitfalls. The ancient Greek dictum of “Know thyself!” could well be
expanded to say, “Know thyself... and thy subject!”
80 << STATISTICS FOR THE SOCIAL SCIENCES

| EXERCISES
Exercise 3.1

Assume that you are developing a written questionnaire. Develop questions and
categories of response (or scoring instructions) that together form the operational
definitions of the following concepts:
1 » Age
2 . Religion
3 . Marital status
4 . Party identity
é . Attitude on environmental problems
6. Attitude on rights of homosexuals
7. Attitude on compulsory national service (military or nonmilitary)
8. Attitude on tolerance toward racial, religious, or linguistic minorities
9. Attitude on tolerance of sexually related publications
10 . Attitude on tolerance of cigarette smoking by others

Exercise 3.2
Assume that you are developing indices in which countries are the units of analysis.
What factors would you consider in developing scales for each of the following?
How might you weight these factors?
Political tolerance
Harshness of criminal penalties
Freedom of religion
Disability awareness
Public safety
Sa
WwW
Ft
GF
OO
= Public health

NOTES
1. As defined in The American Heritage Dictionary of the English Language,
New College Edition (Boston: Houghton Mifflin, 1981), p. 524.
2. Earl Babbie, 7he Practice ofSocial Research, 5th ed. (Belmont, CA: Wadsworth,
1989), pp. 125-6.
= xn ale
ae
| _
-

Tn
S Perien *
2, i: ii) ———
Aa) eee

nen 148 : mia bn am


wae 2 » oe
a += [ay
wy <i nw SS») boeey vee =a evn? 148

: A eorkyeh Tc
at “ Or ;
Fem ji= : ¥ as
>
== (ein ee 16 of 4a ‘alot na © &
a ae A
ear: 7 » s iSteed
ed Oh) ae O =
W KEY CONCEPTS ¥
LASER OETA ERLEOl EE I ENN EERIE BE LE IEE EIRENE LEBEL LESLIE
IIE MIELE DEERE TIBI LEE LILLE LLELL LAL EE
LTE LAAN LABS,

measures of central X-axis symmetric frequency


tendency faxis distribution
mean/arithmetic mean origin (of a graph) positively skewed/
summation symbol frequency polygon skewed to the right
(capital sigma: )~) histogram negatively skewed/
x, y and so on smooth curve skewed to the left
median (Vd.) continuous variable stem and leaf displays
median position (Md. Pos.) unimodal, bimodal, and boxplots/box and
array trimodal frequency whisker plots
cumulative frequency (cf) distributions fractiles
mode modality quartiles
modal class/modal skewness deciles
category symmetry percentiles
[ASR BERD PUSUSS SA UT SC Sta NSBR SED
TE os ieee LEE MLE I AN NS
CHAPTER

Measuring
Central Tendency

¥Y PROLOGUE ¥

Suppose I have now developed my bigotry index and I want to apply it. A
score will be assigned to each person studied such that the higher the score,
the greater that person’s level of bigotry as defined by me. Assume I want to
study two different groups of people, one of which I suspect is more bigoted
than the other. How may I compare the groups to verify my assumption?
One way is to determine for each group a score that reflects the middle level
of bigotry and compare the two to find out who are the bigger bigots. The
score representing the middle level of each group we call an average or,
more elegantly, a measure of central tendency.
BR AAD LSE ELLEN RESELLE NOE INE EEL ES EEE EE NIN EEE SN NIE EEE EELLLL

83
84 @ STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
In this and the next chapter, we examine how to describe a set of scores on
one particular variable for some group so that we may compare that group
to other groups measured on the same variable. Two types of measures exist
for this task: measures of central tendency and measures of dispersion. In
this chapter, we discuss the measures of central tendency (also called
averages or measures of location), which find a single number that reflects
the middle of the distribution of scores—the “average” (meaning typical)
score for that group. We will discuss measures of dispersion in Chapter 5,
and then, with these topics discussed, we will be able to return to the issue
of relationships between variables.

Measures of central tendency Averages or measures of location that find a single


number that reflects the middle of the distribution of scores—the “average” score
for that group.

Measures of dispersion Measures concerning the degree that the scores under
study are dispersed or spread around the mean.

CENTRAL TENDENCY

Suppose you want to study public opinion on the issue of censorship ofthe
arts, specifically, whether governmental agencies funding artists should
refuse to fund erotic or other controversial art projects. Your subjects are
alumni at a 5-year college class reunion. You determine each subject’s
major field of study in college and ask each subject to self-assign a score on
a0 to 10 scale, where 10 indicates the most support for artistic freedom (or
the least amount of censorship). Because this is to be a pretest of a much
wider study, other questions will also be asked. You hypothesize that
alumni who majored in the liberal arts disciplines would be far more in
favor of artistic freedom than those majoring in other fields such as the
sciences, business, or health. Let us assume for computational ease that
your study includes 9 non-liberal arts majors (Group A) and 10 liberal arts
majors (Group B).
The traditional way of seeing whether or not the two groups differ is to
compare the average artistic freedom score for each of the groups. To find
the “average” score (as you probably learned it), we add up all the scores for
each group and divide by the number of students in the group. Suppose the
scores are as follows:
Measuring Central Tendency » 85

Group A Group B
oe) \

Dow
INDDWIN
Xe)
Tie.)
eo
12)
I)
|)
SS
ONG
ON

Total 75

In Group A, we note that we have 9 scores. Adding the 9 scores together, we


get 63. Dividing 63 by 9, we get an average score of7.0. In Group B, we have
10 scores, and the sum of those scores is 75. Dividing 75 by 10, we get an
average score of 7.5. Thus, Group B scored higher than Group A.
Making such a comparison would be simple were it not for the fact that
in reality, there are several kinds of “averages.” The average we found above,
called the mean, is only one of many measures of central tendency, each
appropriate to certain kinds of data or certain analytical needs. In this chapter,
we deal with three of these measures: the mean, the median, and the mode.

THE MEAN

What we have called the “average” (a term we will now avoid since there are
several “averages”) is actually called the arithmetic mean. We will simply call it
the mean since although there are other kinds of means (the geometric mean
and the harmonic mean), only the arithmetic mean will be used in this text.

Arithmetic mean What most people learn in school as “the average.” A measure
of central tendency taking into account the distances from it of all the scores.

Let us label our variable, the artistic freedom score, variable x. We use x
simply to distinguish our variable from other variables that could apply to
the same group. If we also wanted to know the mean social status for the
same group, we could designate status as variable y. We might also want to
know the mean income, and we could designate income as variable z. Right
now, assume that we are interested in only one variable—artistic freedom—
variable x. To find the mean, we add up or sum all the artistic freedom
scores. We ¢all this the summation of x and designate it with the uppercase
86 @ STATISTICS FOR THE SOCIAL SCIENCES

Greek letter sigma (>), known as the summation symbol, followed by


the letter x, which designates what variable we are summing up.

Summation symbol Symbol represented by the uppercase Greek letter sigma (}°).

ix = the summation of x

We will divide } \x by the total number of people in the group, which we


designate as 7, which stands for number, meaning the number of cases
(people, places, or things being measured).

m = the number of cases

The quotient when we divide }°x by 7 is the mean score for our group
along variable x. We designate that as x,often read as x-bar because of the
bar over the x. (Logically, then, the mean social status would by y and the
mean income would be Z.)

x-bar (x) The mean value of the variable x.

x = the arithmetic mean for variable x

Therefore,
e ie
Xe
n

for Group A, x= ya ag 63 =
n 9

Samet
for Group B, ¥ = BoB ol Mae Fis
n 10

Sometimes, the scores for the group are not individually listed but
rather are presented in an ungrouped frequency distribution. Then we
would have two columns of information: column x, which lists every theo-
retically possible score that actually was found in that group, and column f
which lists the frequency or number of times that score actually occurred in
the group. For example, for Group A, we note that only the scores of 8, 7,
and 6 occurred. Thus, the frequency distribution would look like this:
Measuring Central Tendency » 87

Group A
x= e
8 3
Z 3
6 S
In this instance, it is tzappropriate to use the formula for finding the mean
that we used above! Do not add the x column! Do not count up the numbers
in the x column and call it 72/ In this case, 7 is the summation of the f colummn,
and that becomes the denominator for the mean’s formula. Adding the x col-
umn gives meaningless information—we must instead count every the actual
number of times it occurs in Group A. To do this, we multiply each score by
the number of times it occurs, that is, by the frequency in thef column. In
doing this, we generate a new column labeled fx (for f times x). The sum of
that new column, Xfx, becomes the numerator in the formula for the mean.

2 18
3 ge eS
3 | il
2 12
—7=Tf= | Yeas
Note that using this formula with the frequency distribution yields the
same mean as using the original formula for individual data. Conceptually,
the formulas do the same thing, one working from a listing of all scores and
the other from the shorterfx summary data.
The mean is the most mathematically sophisticated and most com-
monly used measure of central tendency of those presented in this chapter.
Mathematically, the arithmetic mean is the value of x that satisfies the fol-
lowing algebraic expression:

20

If we subtract the mean from each of the original scores and then sum the
differences algebraically, then the sum of those differences is 0. Note that
88 << STATISTICS FOR THE SOCIAL SCIENCES

this formula takes into consideration not only the value of each score but
also its distance from the mean (x—*X). Other measures of central tendency
are less sophisticated in that they do not incorporate such distances.

BOX 4.1

More on the Summation Symbol

Those of you who have taken several mathematics courses may be aware
of the fact that we are using }> to say “add up all the scores.” Various
notations placed around )° can be used to exclude certain scores from
the addition. Since in this book, we will have no need to exclude any
scores in a listing, we merely use }) unadorned by other symbols.
Technically, though, the full formula for the mean looks like this:
n

pee
a
1 Xp XQ HZ + Xp
Ca. i
Letting 7 indicate the particular person whose score is being counted,
the numerator is read “the summation of x-sub-7 as 7 ranges from one
(the first person) to 7 (the last person).” Suppose we listed from high-
est to lowest the scores for Group A, associating each person with a
number (7) as if the 7 were his or her name.

Group A
eae eee | n
ee
8 8 |
ne
x= i

i
a Pel |
7 ee sig
l gi
ayOo iat
ve) ln had

sh Sap :
4 kaa _ 6+64+64+74+74+74848+8
3 ome - :

ros Sige fnew


a 6 |

ae
For Group A, we simplify the notation,

fa n
Bs
9
ons
Measuring Central Tendency » 89

BOX 4.2
Making Use of the Definition of the Mean

The definition of the mean provides us with one method of checking


to see if we correctly calculated the mean. Let us go back to the
original scores for Group A since the }\(«—-xX) = 0 formula applies
to individual ungrouped data not in a frequency distribution.
Remembering that the mean for Group A was 7.0, subtract 7.0 from
each value of x and add those differences algebraically.

Group A

X= — DONG =

8 | 7 +1
8 y | +3
8 | 7 +1
a 7 0
7 | of 0
7 i, 0
6 | 7 =)
6 7 =| s
ao | 7 zl
yx = 63 ps Clie 0)

Ke
X= S- = — =70
‘i n ?

Had }*(@-—x) not been 0, there would be a great likelihood that we


had made a mistake in the original calculation of X. If }) @—x) is not 0
but is quite small, however, that may be due to rounding error.
For data in a frequency distribution, the mean is defined by the
following formula:

ii@axil=0
.
Here we factor in the frequency in which each value of x appears
(Continued)
90 @ STATISTICS FOR THE SOCIAL SCIENCES

(Continued)
Group A
= i= | X = xX = aX) f=
8 5 | 7 +1 1x3=+3
7 3 | 7 0 Oxs=70
6 3 | 7 -1 SIS SSB
SI@-Afl= 0
We will encounter such “deviation scores” as }*(~—X) again in the next
chapter.

THE MEDIAN

A second measure of central tendency is the median. In a list where values of


x are arranged from highest to lowest score, the score that falls in the middle
is the median score. Thus, the median is that value of x such that there are as
many scores greater than the median as there are scores less than the median.

Median A value in which there are as many scores greater than the median as
there are scores less than the median.

To find the median, we begin by finding the person, place, or thing that
possesses the median score. This middle position is known as the median
position. Whatever the score possessed by the person, place, or thing at
the median position is the median itself. Note that the median is mot the
median position! The median is the value of the variable that is associated
with the person, place, or thing in the median position.

Median position The person, place, or thing that possesses the median score or
middle position.

Let us take it step by step. First, place the scores in an array, a listing
from highest to lowest (or lowest to highest). Second, find the median posi-
tion (Md. Pos.) by using the following formula:

1
Md. Pos. = ie

Third, find the score associated with the median position. That score is the
median (Md. ).
Measuring Central Tendency » 91

Array A listing from highest to lowest (or lowest to highest).

Now let us look at a specific example, the list for Group A.

ON
GNGN
Coco
Co
SII

The scores are listed in an array from highest to lowest. Since there are
nine scores, 7 = 9, the median position is

ee ees ele
Md. Pos 5 5 5 5

Thus, the fifth person in the array is the one possessing the median score.

Md. Pos. X

©PNW
D~I
KRU ©NADA
YN

ing the
The 7 = column is a convention for identifying the person possess
with the
adjacent value of x. Instead of using a name, we identify the person
as i=2, and
lowest score as 7 = 1, the person with the second lowest score
92 @ STATISTICS
FOR THE SOCIAL SCIENCES

work up to the person with the highest score, 7 = 7. The person at the
median position here is therefore 7 = 5.
Note that we could have counted down from the top, the highest value
of x, and arrived at the same conclusion.

i= x=
1 8
2 8
2S 8 Counting down to the fifth
4 ve person (i = 5), we see that
Md. Pos. 5 7 —— Md. the adjacent value of x is 7.
6 ‘i The median is 7.
rf 6
8 6
) 6

When we calculate the median for Group B, we encounter a new


complication.

Mae 7 Ua De oe
Md. Pos = = 5 me Saal
aniS 2 a

)
8 Whenever the original 72 is an odd
8 number, as in Group A, 7 + 1 is even, and
ERR eee 8 the median position is a whole number.
-
7 Whenever the original 72 is an even
7 number, as in Group B, 7 + I is odd, and
6 the median position is a number with a
WwW
PNM
©
ON
KU 6 .5 decimal (in this case, 5.5).

There is no 5.5th person in the array, so we take the score of the person
just below the hypothetical 5.5th and the score of the person just above that
5.5th score. In other words, we take the score of 7=5 and the score of 7 = 6,
which are 7 and 8, respectively. The median is the midpoint of those two
values and is calculated in the same way that a mean is calculated, by adding
the two scores and dividing by 2. Therefore,
Measuring Central Tendency » 93

GroupB
— i
10 2
2) y
8 8
# 8

Mad. Pos Meort


.= 5.5 ——
6
5
8
5 ER
7+8
se
15
arm

4 7
3 i
iy 6
1 6

Note that in both groups, the means equalled their respective medians.

For Group A, Xx =Md. =7


For Group B, xX=Md. =7.5

However, this is often mot the case. Consider the following problem:

jes x= There is one large score (x = 500)


5 500 plus four much smaller scores.
4 50
Md. Pos. 3. 30——Md. Se eh ees
2 20 n 5
J 10 id ees pent laa
2 2 Z

Where 7 = 3 (the median position),


the adjacent value of x is 30.
Thus, Md. = 30.
Md. #[Link] median is not equal to the
mean.

Thus, we see an important difference between the two measures of


central tendency: The calculation of the mean takes into consideration the
distance between the mean and each score; the calculation of the median does
not. The one extreme value of 500 in our problem causes the mean to increase
in magnitude toward that extreme value, but it does not increase the median.
To illustrate, let us change the previous problem by replacing 500 with 60.
94 @ STATISTICS FOR THE SOCIAL SCIENCES

Ve x= wet S106
5 60 os raha alae

4 50
Md. Pos. 3 30 —— Md. Thus, Md. = 30 and is unchanged from
2 20 the previous problem.
1 10
yx =170 The mean, however, is reduced
considerably:

ses n e.
The mean falls from 122 to 34, whereas the median—not affected by the
actual values of x—remains the same.

Grouped Data

In the past, large data sets were often grouped first to ease the job of
calculating by hand or by mechanical calculator, and then from the grouped
data means and medians were estimated.
Many statistics books present the techniques for doing so, but we will
not cover those techniques for a variety of reasons. First, they are only esti-
mates of the true mean and median, which lessens their value to us. Second,
in the case of the median, the technique is complex and time-consuming.
Third, with today’s calculators and computers, it is possible to find these
measures even for very large data sets.
Why not just use computers to calculate means and medians all the
time and instead of learning the previous techniques? Because without
understanding the logic of the formulas used, a researcher may select an
inappropriate measure for his or her data. Also, the risk of incorrectly inter-
preting the findings would increase. Finally, with small data sets, unless
a personal computer with statistical software is readily at hand, it is faster
to grind out these statistics using a calculator than it is to go to a computer
center, input the data, and wait for a printout.

Finding the Median in a Frequency Distribution

In the case of a frequency distribution, the procedure for finding the


median parallels the procedure for ungrouped data. We find the median
position, remembering that in a frequency distribution, 7 = }*f Therefore,

1 i
Md. Pos. = - = eae
Measuring Central Tendency » 95

We count up or down the frequency column until we encounter the median


position and see what value of x is associated with this frequency,

Group A HED oe oe ae EI ee a
x= = 2 4 He Z
8 Ss
7 3 The fifth pérson in the array is at the median position.
6 ©; Looking at thefcolumn, we see by counting up that
n=y f=9 the first three people have the score of 6, and persons
number 4, 5, and 6 all fall in the category adjacent
tox = 7. Thus, 7 is the median:

This process is sometimes eased, particularly when there are many


scores or 7 is large, by generating a cumulative frequency (cf) column.
We can do this either from the top working down or from the bottom
working up, although we generally work from the bottom up. Working
from the bottom, we begin with the smallest score and enter its frequency
in the cf column.

Cumulative frequency A way of accounting for all the frequencies generated


up to a specific value of x. It is used to determine which value of x is at the
median position.

= cf= Thecfcolumn tells us that in counting our lowest


5 3 score, x = 6, we have accounted for the first three
people @ =1,7=2, and7=3).

We move to the next highest score (x = 7), note its frequency, and add
that to the number in the cf column below it. This number tells us the total
accumulated number of people we have accounted for after passing a score
of x = 7 or below.

f= f= We have now accounted for the first six people


2 6 and 7=6).
@=1,¢=2,17=3,7=4,1=5,
all
SI
oy

Finally, we add to our cumulative frequency column the additional three


people possessing a score of x = 8.
96 @ STATISTICS FOR THE SOCIAL SCIENCES

x= f= cf= Wedo this until we run out of scores possessing


8 3 9 a frequency other than 0. At that point, cf would
7 2 6 bem, and it would remain at that number for any
6 3 3 subsequent values ofx.

Listing all possible scores:

x= f= cf= Weknow already that our median position is 5.


8 a 9 Since cf is 3 where x = 6 and cf is 6 where
vi 3 6 x=/7, we observe that the person possessing
6 3 3 the median position has entered where x = 7.
5 0 0 Persons 7=457 = 5,4nd7 = enter where x = 7.
4 0 0
3 0 0
2 0 0
1 0 0
0 0 0

Since person 7 = 5, possessing the median position, must have entered


where =-7, 7 is the median.

Looking at Group B:

= = = 1 1 10+1 11
ie fi a) Md. Pos. = cane is = = =5.5
Mer aA ele, “ : a
8 3 8
7 3 5 We need to find the scores of persons 7 = 5 and
6 Z 2. 4=0. Whele G = 5.4 = J. Person 2 = 511s Included
among those where x = 7.

We see that since cf for x = 8 is 8, persons 7 = 6, 7 = 7, andi =8 are


entered where x = 8. Thus, person 7 = 6 possesses a score ofx = 8.
Since the score for 7 = 5 is 7 and the score for 7 = 6 is 8, we take the
midpoint of the two to find the median.

74.8 1s The same result is found as when the


Md. = = ai =7.5 median was found from the individual
scores in Group B.

Let us find the mean and median for one more example of a frequency
distribution.
Measuring Central Tendency 97

x= i

10 10
9 20 To find the mean, we will need to determine 7,
8 40 which is \* f. We will need )~ fx, so we construct an
v 20 jx column. Finally, we will need a cf column to help
6 10 us locate the median position and subsequently
5 5 the median itself.
4 5
3) 10
2 30
1) 10
0 5

x= ihe | fx = of=
10 10 100 165
9 20 | 180 155
8 40 320 135
‘i 20 | 140 ”»
6 10 60 iS
» 5 z> 65
4 5 20 | 60
© 10 30 aD
2 30 60 1
1 10 | 10 | 15
0 LS | ew 2
n= )~ f= 165 yi)

Finding the mean:

De- fume E fo =
fs SF ee 24>
165 5.727272 = = 5.727
oe

For the median position:

pei sl 166 =o
es feito
Md. Pos. = vs 2
2 Z a

Looking up the cf column, we see that when we account for the scores
for the
through x = 6, 75 subjects are accounted for; when we account
the subject at the
scores through x = 7, 95 subjects are accounted for. Thus,
Md. = 7.
median position, 7 = 83, enters where x = 7. Therefore,
98 @ STATISTICS FOR THE SOCIAL SCIENCES

USING CENTRAL TENDENCY

Remember that the primary purpose in calculating a measure of central


tendency is to compare it to similar measures. Which class did better on a stan-
dard quiz, section one, section two, or section three? What region of the United
States has the highest voter turnout, the Northeast, the South, or the Midwest?
Are families in Asia larger than families in Latin America? What European
country has the largest per capita gross national product (GNP)? In making
such comparisons, we always compare means with means or medians with
medians; never is the mean of one group compared to the median of another.
The same point holds true for the mode, which will be discussed shortly.
Also, each measure of central tendency assumes a different level of
measurement. The mean requires interval-level data, the median requires
at least ordinal-level data, and the mode is the only one of the three that may
have some limited applicability to nominal data. Accordingly, we may find and
use all three measures on interval-level data; we may use the median and
mode (but not the mean) on ordinal-level data. Only the mode can be used at
the nominal level. Further discussion of this point will come a bit later in this
chapter.
Before going on to the mode, the last of the measures of central
tendency, it would be good to get some practice in both calculating and
comparing these measures by completing Exercises 4.1 to 4.4 at the end
of this chapter. When your calculations are done, compare them to the
answers given at the back of the book.

THE MODE

The mode is a third measure of central tendency. It is a category of a


variable that contains more cases than can be found in either category adja-
cent to it. Generally, it simply is the category with the largest frequency.
Consider the following age distribution:

Mode a category of a variable that contains more cases than can be found in either
category adjacent to it.

Age i=
40-59 15
20-39 30
0-19 10
Total 5D
Measuring Central Tendency » 99

We would call the 20-39 age group the modal class or modal category
since it has a higher frequency than either adjacent category. Note that like
the median, but unlike the mean, extreme values of the variable have no
impact on the value of the mode. One unusual characteristic of the mode
is that there may be more than one mode in a particular frequency distri-
bution. For example,

Age iis
50-59 15
40-49 45
30-39 20
20-29 10
10-19 35
0-9 5
Total 130

Both the 40-49 and the 10-19 age groups have more cases than the adjacent
classes (above or below them), and thus both are modal classes. Note that
they need not each have the same frequency.

Modal class or modal category Where data have been grouped, a class interval
or category that contains more cases than can be found in either category
adjacent to it.

Another characteristic of the mode is that it requires a fairly large 7 for


all modes to be correctly identified. Suppose for a sample of 26 respon-
dents, the following age distribution had been observed:

Age =
50-59 3
40-49 9
30-39 4 Modal classes are 40-49 and 10-19.
20-29 2
10-19 7
0-9 1
Total 26

If the sample size were increased from 26 to 130, the pattern with two
modes might appear as before:
100 << STATISTICS FOR THE SOCIAL SCIENCES

Age i
50-59 15
40-49 45 The same two modal classes
ehlrew, 20 observed before, 40-49 and 10-19, appear.
20-29 10
10-19 5D
0-9 oe
Total 130

But it is also possible that the identity of one of the modes could disappear
with increased sample size.

Age is
50-59 1S
40-49 45 When the sample size was 26, the 10-19 class
30-39 20 with f= 7 appeared to be a mode, but when
20-29 20 the sample size increases to 130, it becomes
10-19 1S) clear that the 10-19 “modal class” was really
0-9 15 due to the small size of the original sample.
Total 130

For individual, interval-level data, we can most clearly understand the


mode by changing the data format to an ungrouped frequency distribution
and then generating a graph of that distribution. To graph the frequency dis-
tribution, we lay out two perpendicular lines, a horizontal line labeled x and
a vertical line labeled f referred to as the x-axis and the f-axis, respectively. '
These axes each have numerical scales much like those on rulers.

x-axis and f-axis Two perpendicular lines, a horizontal line labeled x and a vertical
line labeled f.

Where the x-axis intercepts the faxis—the origin of the graph—both x


and fare zero. The numbers on the x-axis increase as one moves to the right
of the origin. The distance between each unit and the unit that follows is
always the same; that is, the distance from zero to 1 is the same as the dis-
tance from 1 to 2, and so on. On the f-axis, the units increase in size as one
moves above the origin. Distances between the units on the f-axis are also
equal, although these distances need not be the same as the equivalent
distances on the x-axis.

Origin Point where the x-axis intersects the faxis.


sateen
c eeeececcec
Measuring Central Tendency » 10]

Let us graph the ungrouped frequency distribution from the section


“Finding the Median in a Frequency Distribution,” for which we found
the
mean of 5.727 and the median of 7.

x= fis

10
20
40

Nn

10
30
10
to
Wi
PS
SS
INS)
my
SoS)
(C9)
| 5
The numbers on the axes reflect the ranges of scores. Since in our problem,
x ranges from 0 to 10, we lay out units of 0, 1, 2, 3, and so on, up to 10. Since
f ranges from 0 to 40, we lay out distances on the f-axis in units of 5: 0, 5, 10,
15,..., 40. This procedure is illustrated in Figure 4.1. For each value of x,
we find its corresponding value of fand move up the graph above the value
of x, placing a dot at the point where we are adjacent to the appropriate
f value. Thus, since where x = 0, f= 5, we move up the f-axis to f= 5 and
place a dot. Since where x = 1, f= 10, we move up directly above x = 1 until
we are on the same level as f= 10 and place a dot. We do this until we
exhaust all values ofx in our range of scores.
The connection of the dots may be done in three general forms. The
first of these, a frequency polygon, is formed by drawing a straight line
from each dot to the next dot, as x increases. The end points would be on
the x-axis one-half unit above the highest appearing frequency and one-half
unit below the lowest appearing frequency. The second form is a type of bar
graph known as a histogram. Bars are created from one-half unit below
each value of x to one-half unit above that value. The third form is made by
joining the points in a smooth curve.

Frequency polygon Connection of dots formed by drawing a straight line from


each dot to the next dot, as x increases.
Histogram Graph in which bars are created from one-half unit below each value
of x to one-half unit above that value. The Faxis indicates the frequency of each
score’s occurrence.
Smooth curve Connection of dots similar to a frequency polygon but generated
by a curved line fitting through the dots instead of a series of straight lines.
102 STATISTICS FOR THE SOCIAL SCIENCES

Technically, a smooth curve is most appropriate where 77 is large and x is


a continuous variable. By a continuous variable, we mean one that is not
limited to whole number scores. For example, suppose 1000 people were
measured on a quiz containing 100 questions worth 1/10 of a point each. While
scores range as before from 0 to 10, scores such as 9.7 or 6.9 could also appear.
The smooth curve would give an accurate portrayal of the pattern produced
by the many resulting dots. In the case of a smaller 7 or where x is a discrete,
rather than a continuous, variable (the scores are only whole numbers from 0
to 10), the use of a smooth curve seems less logical than the other forms of
graphing. Nevertheless, the smooth curve is often used, no matter what the 7
or the nature of the variable. We use it in our example because although the
scores only range from 0 to 10 in whole numbers, they represent an underly-
ing continuous variable. It is similar to the age variable. Although we generally
express adult ages only in years, underlying the years are months, weeks, and
even smaller units of time, which we do not express but exist nonetheless.

Figure 4.1. Graphs for the Ungrouped Frequency Distribution of the “Finding
the Median in a Frequency Distribution” Section

SSS x
I) eae Bere ts SIG es Ab 6e 7 SOO
Plotting the dots A frequency polygon

f f

40 ral
35 354
30 50S
7255) Do)|
20 204
15 - 55]
10 10 a I, We)|

i.
Boe Lo pa F abe x
5
O Sau Sn See See ee | xX

Be eS AN sy. Kon 70 teh 18h ILO) 2 SAGs


7 78 29) 110
A histogram (bar graph) A smooth curve
Measuring Central Tendency » 103

‘seen SNES tS SS ESE

Continuous variable A variable that is not limited to a finite number of scores.

Our obtained results for all three measures of central tendency are
presented in Figure 4.2. The modes are those values of xwhere the curve
peaks, in this case, where x = 2 and again where x = 8. (Remember that
the modes are the values of x and not their respective frequencies. It
would be wrong to say that the modes are 30 and 40.) Since there are two
modes, we say that the distribution is bimodal. If there were only one
mode, it would be called unimodal; if there were three modes, the dis-
tribution would be trimodal, and so on. By the term modality, we mean
the number of modes found in the frequency distribution. A great deal of
information about a frequency distribution can be communicated verbally
just by indicating its modality and skewness, another characteristic to be
discussed shortly.

Unimodal, bimodal, and trimodal A distribution with one, two, and three modes,
respectively.

Modality The number of modes found in the frequency distribution.

Figure 4.2 | Measures of Central Tendency


104 << STATISTICS FOR THE SOCIAL SCIENCES

INTERPRETING GRAPHS

Although graphs have many kinds of applications, as they are being used
here, graphs are pictures of frequency distributions. It may take a while to
get used to them, but once you have become familiar with how to read the
graphs, you will appreciate that sometimes a picture really is worth a thou-
sand words, give or take. Examine Figure 4.2. Remember that the x-axis
shows the range of all the scores under study—in this case, 0 through 10.
The f-axis shows the number of people possessing each score listed on the
x-axis. The height of the curve at any given value of x is the number of
people having that value of x (i.e., sharing the same score).
If we start at the origin of the graph in Figure 4.2, we see that five people
share a score [Link] we move to the right along the x-axis, we see that more
people have scores of 1 than of 0. The curve rises from a frequency of 5 to
a frequency of 10 and then continues rising until, atx = 2, 30 people possess
that score. At this point, we see that as the scores rise, so do the number of
people possessing the score. Then, however, the curve begins to drop to 10
and then, atx =4 andx =5, to 5. The picturing in the graph so far is a steep
hill, rising until it peaks at the mode of x = 2 and then falling offas the scores
continue to increase. The pattern shows a clustering around the score x = 2.
If our frequency distribution were unimodal, the frequencies of scores
to the right of x= 5 would continue to diminish. We would conclude that
most people were scoring at or near the mode of 2. The area under the
curve on our graph corresponds to the number of people with each score
or the number of people within a region of scores. In this instance, 65
people have scores in the region of x = 0 tox = 5, with their scores cluster-
ing around the mode ofx = 2.
As we continue along the x-axis past x = 5, however, the curve does not
drop off; it begins rising again. It rises as x increases, until it reaches another
mode at x = 8, where it maximizes and then begins to decline. The bimodal
nature of the distribution hints that we may be identifying two different
groups of people, a group whose scores cluster around 2 and another group
whose scores cluster around 8.
Imagine that Figure 4.2 represented the scores on the first quiz in a
course in the French language—say, French 102, the second course in the
sequence beginning with French 101. Not surprisingly, French 101 is a pre-
requisite to French 102, but suppose nobody checked for prerequisites, so
that people could register for French 102 without French 101. Maybe 65
students had taken some high school French, decided they didn’t need to
take French 101, and so signed up for French 102 as their first college-level
course in that language. But their instructor, Professor Javert, decides to
make his students’ lives miserable with a very tough 10-question quiz. The
Measuring Central Tendency 105

students who had taken French 101 end up with scores clustering around 8;
the ones without the prerequisite have scores clustering around 2 and find
themselves in the proverbial sewers of Paris.
Note also the relative height of our two peaks. Since the peak on the
right is higher than the peak on the left, we may conclude that more people
have scores clustering about the mode of 8 than the mode of 2. Had the left-
hand peak been higher, we would conclude the opposite: More people were
clustering around 2 than 8. Luckily for our French students, more people
had taken the prerequisite than had not.

CENTRAL TENDENCY AND LEVELS OF MEASUREMENT

Now that all three measures of central tendency have been presented, we
note again that the usage of a measure of central tendency is determined in
part by the level of measurement of the data. The mode is the only measure
that may be used on data of all measurement levels. We have already seen it
applied to interval-level data, both ungrouped (where we found modes of 2
and 8) and grouped (where 40-49 was the modal class). Similarly, we could
apply the mode to ordinal-level data.

Feelings of Verbal Efficacy fie


Very High 15
Somewhat High 30
Moderate 50
Somewhat Low 15
Very Low _10
Total 120

The modal category of feelings of verbal efficacy (effective verbal communi-


cation) is the moderate category.
We could also apply it to nominal-level data, although care must be taken
to remember that a// we learn is which category had the largest frequency.

Region of Canada pe
Atlantic Canada 10
Quebec 30
Ontario eB)
Prairie Provinces 20
British Columbia iS
Territories S
- Total A
106 << STATISTICS FOR THE SOCIAL SCIENCES

So, for our 115 respondents, the modal Canadian was from Ontario. (You
could say that the modal Canadian was from either Quebec or Ontario; how-
ever, with no ordering to our categories—these are nominal data—how big
a frequency do you need to determine that it is a mode?)
By contrast, the median assumes either ordinal- or interval-level data, so
we can find no median region of Canada. We have calculated the median for
interval-level data. To calculate it for ordinal-level data, let us reexamine our
Feelings of Verbal Efficacy example.

Feelings of Verbal Efficacy f= =

Very High 15 120


Somewhat High 30 105
Moderate 50 ie
Somewhat Low 15 25
Very Low a0) 10
Total 120

n+1_ 12041 121


Md. Pos. = — =60.5
a Zz

Examining the cf column, we see that both the 60th and the 6lst respon-
dents enter at the moderate category. Thus, moderate is the median level of
Feelings of Verbal Efficacy.
Of the three measures of central tendency, the most sophisticated
measure—the mean—is reserved for the most sophisticated levels of
measurement—interval level and ratio level. The mean assumes that one is
working with scores whose true distances from the mean can be ascer-
tained. Since rankings do not reveal actual distances, they should not be
used for finding a mean.

SKEWNESS

We can also describe the nature of a frequency distribution by making


reference to skewness, the extent to which the distribution deviates from
symmetry or the balance between the right and left halves of the curve.
In fact, we will see that we determine skewness from two measures of
central tendency—the mean and median. From skewness plus the number
of modes, we can visualize a very close replica of the actual appearance of
a given frequency distribution, A distribution with no skewness (skewness
equals zero) is said to be a symmetric distribution. In a symmetric
Measuring Central Tendency 107

distribution, the curve to the left of the mean is a mirror image of the curve
to the right of the mean.

Skewness The extent to which the frequency distribution deviates from symmetry.

Symmetry The balance between the right and left halves of the curve.

Symmetric frequency distribution ~ A frequency distribution with no skewness


(skewness equals zero).

An illustration is given in Figure 4.3. When a curve is unimodal and


symmetric, the mean, median, and mode are all of the same value, as illus-
trated by the vertical line in the center of the figure. Imagine Figure 4.3 lying
flat on a page of tracing paper folded down the vertical line at the point of
central tendency. If we folded the page as we would close a book, the curve
on the side to the left of the fold would fall exactly on the curve to the right,
as shown in Figure 4.4. The left side is a mirror image of the right side; thus
we have symmetry.
In a bimodal, symmetric distribution, as shown in Figure 4.5, the modes
are different from the mean or median, but the mean and median are equal.
Symmetry is illustrated in Figure 4.6, where the page is folded along the
vertical line where x = Md., and the curve to the left of the fold falls directly
over the curve to the right.
Ifa distribution is mot symmetric, we say it is asymmetric or skewed. For
example, look at Figure 4.7. Here the far right (called the right tail) of the

Figure 4.3. A Unimodal, Symmetric Frequency Distribution


108 < STATISTICS FOR THE SOCIAL SCIENCES

Figure 4.4 The Result When the Left Side of the Curve Falls Exactly on the
Right Side—Symmetry

4 fold

x
ve

Figure 4.5. A Bimodal, Symmetric Frequency Distribution

Figure 4.6 Symmetry Illustrated by Folding the Left Side of the Distribution in
Figure 4.5 Over the Right Side

fold
Measuring Central Tendency » 109

Figure 4.7. An Asymmetric or Skewed Distribution

|
|
| |
| |
| |
I |
| |
| |
| |

t Mo. Md. x |
left tail right tail

curve extends much farther than the tail on the left.’ If we fold the page at
the mean, as in Figure 4.8, the left side of the curve does not fall directly on
the right side. There is mo mirror image.
When a curve is skewed, we indicate the direction of skewness with
reference to the longer tail. Thus, in Figure 4.7, the distribution is skewed
to the right or, more formally, positively skewed, since skewness is in the
direction of increasing positive values on the x-axis. If the left tail were
longer than the right tail, then the curve would be skewed to the left or,
more formally, negatively skewed, since in that case, skewness would run
in the direction of increasing negative values on the x-axis. Both cases are
illustrated in Figure 4.9.

Skewed to the right or positively skewed Skewness is in the direction of increasing


positive values on the x-axis.

Skewed to the left or negatively skewed =Skewness runs in the direction of


increasing negative values on the x-axis.

In both examples in Figure 4.9, note that in comparing the mean to the
median, the mean is always the measure of central tendency pulled most in
the direction of skewness, the direction of the more extreme values of x. We
already alluded to this phenomenon in our discussion of the median, and it
points to a situation where we must choose between the mean and median
as the most appropriate measure of central tendency for describing a
particular distribution. (Ignore the mode for the moment.) Generally, the
arithmetic mean, being the more mathematically sophisticated measure of
the two, is preferable to the median. In fact, the mean is so widely used as a
110 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 4.8 — Result of Folding Figure 4.7 Over Right at the Mean

told 33 —=—0

¢ \ left side
‘ of curve

*<| right side of curve

Figure 4.9 Examples of Skewed Distributions

Mo. Md. x
Positively Skewed

x Md. Mo.
Negatively Skewed

measure of central tendency that, as we know, most people call it the aver-
age. Nevertheless, if a distribution is highly skewed, the median may be
more appropriate or honest than the mean for explaining central tendency.
A clear example is found in reporting information pertaining to the
variable income. In 1995, for instance, the median per capita income for
the United States was around $22,000. Here the median is preferable to the
mean because although the dollar range below the median is $22,000, the
Measuring Central Tendency » sia)

same range above the median ends at $44,000. Many people earn incomes
above $44,000 and have the impact of pulling the mean higher in the direc-
tion of skewness. Thus, the mean figure is much higher than the $22,000
median income. Think of it this way: In calculating the median, a person
earning $1 million per year has the same impact or weight as the person
making $1.00 per year. In calculating the mean, however, it could take one
million individuals earning $1.00 per year to counterbalance the impact of
one person earning $1 million. So, although the mean is usually the mea-
sure of choice, we calculate the median along with the mean to add a degree
of protection to our procedure by screening out those instances where the
mean may mislead.
In Figure 4.10, symmetry and skewness are compared for both unimodal
and bimodal distributions, and points of central tendency are noted for each
example. Be aware that the two skewed bimodal distributions are drawn so
that the lower peak lies in the direction of skewness, but this does not have
to be the case. It is sometimes very difficult to ascertain skewness on bimodal
distributions; the examples in Figure 4.10 are for illustrative purposes.

Figure 4.10 Modality and Skewness of Frequency Distributions With Reference to


Central Tendency

Unimodal
Negatively Skewed Symmetric Positively Skewed

>

x Md. Mo. xX Mo. Md.x


= Md.
Mo.

Bimodal
Negatively Skewed Symmetric Positively Skewed
f ~~
112. < STATISTICS FOR THE SOCIAL SCIENCES

OTHER GRAPHIC REPRESENTATIONS

Stem and Leaf Displays

In the past several years, a graphic relative of the histogram, known as the
stem and leaf display, has enjoyed growing popularity as a means of summa-
rizing social data. A stem and leaf display combines the visual effect of a his-
togram but preserves the actual scores in small- to medium-sized data sets.

Stem and leaf display A graphic representation that combines the visual effect of a
histogram but preserves the actual scores in small- to medium-sized data sets.

Note the following frequency distribution, which appeared earlier in this


chapter.
Age f=
50-59 3
40-49 v,
30-3? +
20-29 2
10-19 P
0-9 =f
Total 26

Suppose we list, in ascending order, the actual ages of the 26 people


whose ages constitute the frequency distribution: 8, 10, 12, 13, 16, 17, 17,
18, 21, 26, 32, 33, 35, 38, 40, 41, 42, 43, 44, 45, 47, 47, 48, 52, 53, 59.
To get our stem and leaf display, we list our class intervals in descending
order from the top. Each class interval becomes a “stem” on the stem and
leaf display. We then go through the 26 scores just presented, find the
score’s appropriate class interval, and list the last digit of the score to the
right of the appropriate class interval. For instance, the first score in the list-
ing is 8. We place it beside its class interval, the 0 to 9 stem, and it becomes
the “leaf.” The second score on our listing, 10, is in the 10 to 19 class inter-
val. Thus, the last digit of that score, 0, becomes the first leaf on the 10 to 19
stem. We get the following results for our 26 scores:

Age
0-9 | 8
LO=19 | O21 Tigas
20-29 | 16
30-39 | 23518
40-49 | 012345778
50-59 | 239
Measuring Central Tendency p> fles}

Sometimes, instead of listing the whole class interval as


a stem, we use
only the first digit of the lowest end of the class interval, as follows:

| 8
| 012:3°67
78
| 16
| 2358
| ON Besos
RF
WNM
AR
© | 239
Also, while the scores were ordered from lowest to highest in our
listing, that is not necessarily a requirement for our display. Suppose we had
the following ages, listed in no particular order: 47, 25, 42.50, 58,37, 3521,
42, 42, 45, 51, 32, 48, 36. The stem and leaf display for these 15 ages would
be as follows:

Age
20-29 pul
30-39 hoe oO
40-49 IE Pate
50-59 681

In Figure 4.11, we see how this last stem and leaf display resembles a
histogram. In that figure, the histogram is drawn with horizontal rather than
vertical bars. The bars resemble the display of leaves on each stem but, of
course, the bars do not tell you the original scores, whereas the stem and
leaf display does.

Boxplots

Boxplots or, as they are often called, box and whisker plots, are
useful representations of small data sets for which the kinds of graphs
presented previously would not generally provide useful information. For
example, suppose we took the 10 youngest ages from the first example used
im tne previous section. [hese are 8, 10,12, 13, 16, 17, 17, 18,21, 26.

Boxplots or box and whisker plots Useful representations of small data sets, where
more traditional graphs cannot be used effectively.

To understand boxplots, we must first introduce the concept of fractiles.


Fractiles divide scores into smaller groups of scores of approximately equal
size. We have already learned one such fractile, the median, which divides the
114. STATISTICS FOR THE SOCIAL SCIENCES

Figure 4.11 A Stem and Leaf Display and Equivalent Histogram

Age
20-29 5 1
30-39 H 8
40-49 7 2 Z 5 8
50-59 5 8

Age
20-29
30-39
40-49
50-59

scores into two subsets: the half of the scores less than the median and
the half greater than the median. We may also divide our scores into 4 sets
of scores, called quartiles; 10 sets of scores, called deciles; or 100 sets of
scores, called percentiles. You may already be familiar with these terms in
relation to standardized tests where, in addition to raw scores being
reported, scores are often also reported as fractiles. Thus, if you took such a
test and scored in the 87th percentile, 87% of the scores were below or equal
to your own.

Fractile Measurement that divides scores into smaller groups of scores of


approximately equal size.

Quartile, Decile, Percentile |©A number that divides scores into sets of 4 (quartiles),
10 (deciles), or 100 (percentiles) and indicates for each individual studied the
number of people (or other units of analysis) his or her score exceeds.

Let’s first find the median for our data set. Since 7 = 10, the median posi-
tion is 11/2 or 5.5. Using the procedures already learned, we take the mid-
point of the fifth and sixth scores in the sequence to find the median. Here,
the fifth score is 16, and the sixth score is 17. So the median is 16.5, and there
are 5 scores below 16.5 and 5 scores above it. To find the first quartile—the
lowest one fourth of our scores—we in effect find the “median” of the given
scores below the median for all the data. For these scores—8, 10, 12, 13,
16—the median position is 3 and the median is 12. Thus, we say that the first
quartile is 12, implying that the lowest one fourth of the scores is below 12.
Measuring Central Tendency » {LN

The second quartile for our full data set is really the median of 16.5. This
is because the two lowest quarters of the scores lie below 16.5. (Note that
1/4 + 1/4 = 1/2, so the lowest half of all the scores is the same as the two
lowest quarters of the scores.)
Finally, the third quartile divides the highest one fourth from the second
highest one fourth of the scores. To find it, take the scores greater than the
median of the original data set, 16.5, and find the median of that subset. The
Scores aren 7; 18)2 ane 26: the median position is 3; and the median
is therefore 18.

To summarize:
26 is the highest score
18 is the third quartile
16.5 is the median and second quartile
12 is the first quartile
8 is the lowest score

A simple boxplot of these scores consists of a rectangle (our box) beginning


at the first quartile and extending to the third quartile. The box is subdivided
into two parts by a vertical line at the median or second quartile. From either
end of the box, there is a straight line (a whisker) running to the highest value
and the lowest value, respectively. Our boxplot is depicted in Figure 4.12.
The location of the median line in the box may also tell us the skewness
of the data. If the left “sub-box” is the same size as the right sub-box, the dis-
tribution of scores is symmetric. If the right sub-box is much larger than the
left one, the data are positively skewed, and if the left sub-box is much larger
than the right, the data are negatively skewed. In Figure 4.12 the left sub-box
is slightly larger than the right one, suggesting a negative skew. If we calcu-
late the mean for the data, we find that it is 15.8, less than the median, and
it confirms our suspicion of negative skewness. In some cases, a skewness
decision may also be made based on the length of the whiskers, with the
longer whisker indicating the direction of skewness. However, in Figure 4.12,
although the right (positive whisker) is longer than the left one, we already
know that the skewness is negative, not positive. So we really cannot use
whiskers to determine skewness unless the bigger sub-box and the longer
whisker are both on the same side.
Often, extremely large or extremely small scores, known as outliers,
which are far from the range of the other scores, are left out of a boxplot.
This further confounds our use of sub-boxes and longer whiskers to deter-
mine skewness, so always calculate the mean to make sure. One convention
used to eliminate outliers is to run the whiskers out not to the low and high
116 << STATISTICS FOR THE SOCIAL SCIENCES

scores but only to the 10th percentile on the left and the 90th percentile on
the right.
Suppose that the lowest score in our data set had not been 8 but 1. The
new boxplot is redrawn in Figure 4.13, where both the left whisker and left
sub-box are greater than their right-sided equivalents. Negative skewness is
clearly indicated. The mean for our data (indicated by an arrow) is recalcu-
lated to be 15.1, is less than the median, and confirms negative skewness.
(This figure also suggests that a better name for this kind of plot might be a
“box, whisker, and toothpick” plot!)

Figure 4.12 A Boxplot

CONCLUSION

Measures of central tendency enable us to compare different groups to


determine the relative relationship of their “middle” values. However, we
must be aware that each of the measures discussed has its own idiosyn-
crasies. First, each is keyed to certain levels of measurement. Also, skewness
of the data may make the use of the mean problematic. Finally, if our group’s
size is too small, false modes may appear in a graph.
Yet, knowing central tendency, modality, and skewness enables us to
describe each group we are studying very accurately. We have covered some
very powerful descriptive tools. Another useful set of descriptive tools, mea-
sures of dispersion, will be presented in the next chapter. There we will be
looking for measures that give us an idea of the extent to which scores are
clustered about the central tendency.
Measuring Central Tendency » 117

Chapter 4: Summary of Major Formulas

Individual Data Frequency Distributions


The Mean The Mean
ao SS zu 2Se _ Defx
nN = : nh Dai

The Median Position The Median Position

iteros Mattie 1
aa 1

RE Oe ee xk ae eee |

EXERCISES
Exercise 4.1
Suppose you are interested in studying the impact of oil wealth on those Middle
Eastern countries fortunate enough to be petroleum exporters. To keep calculations
simple, you limit your study to the Arab nations situated on the Asian continent,
leaving out all non-Arab states as well as the Arab states of North Africa. Compare
oil-exporting countries with nonexporting countries in terms of male life expectan-
cies at birth. (In line with patterns elsewhere, female life expectancy is greater than
male life expectancy. In the case of the nations listed below, female life expectancy
is anywhere from 1 to 5 years greater than for males.)

Oil Exporters? Nonexporters

Male Life Male Life


Country Expectancy Country Expectancy

Bahrain 73 Jordan 71
Iraq 66 Lebanon 68
Kuwait io Syria 67
Oman 70 Yemen 58
Qatar ie
Saudia Arabia 69
United Arab Emirates 74

Calculate the means for each group and compare them. What are your
conclusions?
118 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 4.2
Calculate the medians for the data in Exercise 4.1 and compare them. What are
your conclusions?
One possible confounding issue is the case of Iraq, which did export a limited
amount of oil during this period but was under severe restriction during much of
this time as a consequence of the Gulf War and Saddam Hussein’s response to the
international community. Under normal circumstances, it would have exported
more oil and probably had improved life expectancy.

Exercise 4.3
Exclude Iraq from the data of Exercise 4.1 and make a new comparison of both
means and medians. What are your conclusions now?

Exercise 4.4

Let us take most of the Middle East and North African countries and compare oil
exporters to oil importers with regard to the number of people per television set.
Our assumption would be that the exporters would be able to afford more TV sets
and thus have fewer people per television set. Rounded to the nearest whole
number, we get the following. Note that the scores below are placed in ungrouped
frequency distributions; be sure to use the formulas appropriate for such cases.
What are your conclusions?

x = number of people per television set

Oil Exporters? Nonexporters


X= f= <= i=
16 1 36 |
V3 2 16 |
12 1 13 5
10 1 2 1
g 1 9 |
4 1 4 1
3 2 3 |
2 2

Exercise 4,5
In the graphs to the left, the modes are indicated by an M and the other central
tendencies by the letters a through h. Match each of the six distributions on the left
to the appropriate statement: Each of | through VI is used once.
Measuring Central Tendency » 119

I, Bimodal, positively skewed


ll. xX = Md. # Mode(s)
Ill. Unimodal, symmetric
IV. __ Trimodal, negatively skewed
Vo KX < Md. < mode
VI. __ Unimodal, positively skewed

Exercise 4.6
Using the graphs on the left, indicate for
letters a through h whether each is a mean
or a median. (The letter M already indicates
til.
a mode.)

ae
IV. bo ee
©
d.

V.
ae ! eo
Mgh M M {
g.
VI. | None of the above
h.

Exercises 4.7—4.9
In Exercises 4.7 to 4.9, you will be working with the following simulated
data. A sociologist is looking into the relative importance of selected social and
economic problems as perceived by a variety of groups in a particular county.
Respondents are asked to rate the importance of each issue on a scale ranging from
0 (unimportant) to 100 (most important). The issues are as follows:

HEALTH: Health care costs


SUICIDE: Suicides among senior citizens
GROWTH: Economic growth in the county
ABUSE: Child abuse
INSURANCE: Availability of health insurance
AIDS: AIDS and HIV cases in the county

Data were collected from two groups:


Officials: All elected county officials and the major appointed officials
Social Workers: A sample of social workers and related caregiving profession-
als working either directly for the county or for private agencies in the county
120 << STATISTICS FOR THE SOCIAL SCIENCES

Following are the mean importance ratings, by issue, for each of the two groups.

Issue Officials Social Workers


HEALTH 74.95 PO22
SUICIDE 48.59 60.20
GROWTH 03.52 Da.2
ABUSE 62.47 71.30
INSURANCE 73.87 73.82
AIDS 70.00 72.50

To reach conclusions about a given issue, we compare the means of the officials
with those of the social workers. If we subtract the mean of the social workers from
the mean of the officials, we will see the amount of difference between the two
groups. Also, if the difference is positive, then officials consider the issue more
important than do the social workers. Negative differences mean the social work-
ers consider the issue more important than do officials.
In the case of AIDS, for instance, we see that social workers consider that issue
to be slightly more important than do the officials.

Officials Social Workers — Difference


AIDS 70.00 F250 —2.50

Exercise 4.7

1. How are health care costs viewed by the two groups?


2. Which group is more concerned about economic growth?
3. Compare the “social” to the “economic” variables. Which group (predictably)
considers the social issues most important?
4. Do the groups really differ on the issue of health insurance coverage?

Exercise 4.8
Suppose we subdivide the officials into two groups: elected and appointed.

Issue Elected Appointed

HEALTH 78.64 70.20


SUICIDE 42.10 50.00
GROWTH 68.75 62.75
ABUSE 60.00 65.50
INSURANCE 65.25 76.66
AIDS 68.00 72.00

Compare the two groups in terms of the importance of each issue. What overall
conclusions do you reach?
Measuring Central Tendency > 121

Exercise 4.9
Suppose we subdivide the social worker group into two groups: managers and case
workers.

Issue Managers Case Workers


HEALTH 84.20 69.95
SUICIDE ~ 59.00 62 25
GROWTH 57 52 SOA
ABUSE 72.00 70.50
INSURANCE 7362 73.82
AIDS 70.00 74.50

Compare the two groups in terms of the importance of each issue. What do you
conclude?

Exercises 4.10—4.12
In Exercises 4.10 to 4.12, you will be working with data from a hypothetical study
of a major North American automobile manufacturer. The variables, all on scales
ranging from 0 to 100, are as follows:

ATTEND: Employee attendance/participation. The percentage of workdays


(excluding paid vacation) that the employee was not absent in the 12 months
prior to the study.
BOARD: The employee’s support for policies supported by the corporation’s
board of directors in the 12 months prior to the study.
DIV: The employee’s support for positions taken by the director of the
employee’s own division in the 12 months prior to the study.
SECUR: The employee’s feelings of security in keeping his or her present job
during the coming year.
PARTIC: The employee’s support for a participatory decision-making approach
similar to the one used by several Japanese competitors.
OPPOR: The employee’s self-assessment of opportunity for promotion/
advancement in the coming year.
UNION: The employee's support for expanded trade union activities in the firm.
SALARY: The employee's perception of the likelihood of receiving a major pay
raise in his or her current job sometime in the next year.
From this study, two data sets were established.
MGTPOP: All 89 of the highest-level managers in the corporation (top and
upper-middle levels).
EMPLOY: A randomly selected sample of 50 employees working under the top
and middle managers at the firm.
122 << STATISTICS FOR THE SOCIAL SCIENCES

Following are the mean scores for the management population and the employee
sample.

Variable MGTPOP EMPLOY


ATTEND 9735 92.92
BOARD 57.4) 37.38
DIV 74.94 TOL.
SECUR 56.16 50.70
PARTIC 48.59 60.20
OPPOR 42 52 36.64
UNION 62.47 71.30
SALARY 5307 53.82

Exercise 4.10
1. Do managers and employees differ in terms of attendance?
Which group supports the board of directors’ decisions more?
Which group has the greater sense of both security and opportunity?
Rw
Assuming that support for union activity expansion is a measure of employee
discontent, is your conclusion in comparing the means for UNION consistent
with your conclusion in Question 3?

Exercise 4.11
The management group was broken down into top and upper-middle management
subgroups. Mean scores for these subgroups were as follows:

MGTPOP
Top Upper-Middle
Variable Management Management
ATTEND 93.07 91.80
BOARD 609.79 47.22
DIV 70.20 78.64
SECUR LENS 39.80
PARTIC 18.46 7210
OPPOR Poy 16.48
UNION 32.74 85.66
SALARY 76.66 36.10

1. Compare the two groups of managers in terms of


a. Attendance
b. Support for the board of directors
c. Support for division directors’ decisions
2. Do you sense a certain level of discontent in upper-middle management?
Measuring Central Tendency » 123

Exercise 4.12
The employee group was divided into white-collar and blue-collar subgroups. Mean
scores for these subgroups were as follows:

EMPLOY
Variable : White-Collar Blue-Collar
ATTEND 93 31 92.38
BOARD 23.89 56.00
DIV 84.20 69.95
SECUR 3151 77.47
PARTIC 81.89 30.23
OPPOR 1342 69.04
UNION 92.44 42.09
SALARY 35.00 79.80

Compare the two groups along the eight variables. What can you conclude about
them? What conclusion, if any, was unexpected to you?

Exercise 4.13
A social psychologist has developed an index to measure extroversion. It ranges
from 0 to 59 (59 = most extroverted). She administers the index to one of her
classes and obtains the following scores: 7, 8, 9, 10, 12, 14, 17, 20, 21, 22, 24, 28,
28, 29, 30, 30, 33, 34, 36, 37, 38, 39, 41, 43, 45, 47, 48, 52,57, and 59. Prepare
a stem and leaf display of the scores.

Exercise 4.14
For the data in Exercise 4.13, find the median, first quartile, and third quartile.
Prepare a boxplot.

Exercise 4.15
A math anxiety scale ranges from a low score of 0 to a high of 100. This scale was
administered to 35 students in a college orientation program. Following are the
scores: 32,20, 19, 89, 38, 39, 12, 65, 75, 21, 29, 27, 27, 93, 43, 54, 21, 33, 19,
9, 92, 18, 20, 77, 88, 47, 35, 87, 16, 87, 25, 23, 76, 22, 88. Prepare a stem and
leaf display. What can you conclude about the distribution?

Exercise 4.16
Prepare a boxplot for scores above 50 in the math anxiety scale data of Exercise
4.15. How are the scores skewed?
SSE ZI ERR SOC ESE SALSA EE IEEEE LET ERLE EILEEN EELS IIEOSE LEC LIELLELEEEEE IEA, AEE ELLE DE EEE EIN TELE SDS EEE LETTE ELE,
124 << STATISTICS FOR THE SOCIAL SCIENCES

NOTES
1. In later chapters, what we call the f-axis here will be used not for a
frequency but for a second variable, y. At that time, we will refer to that axis
as the y-axis.
2. In Figures 4.7, 4.9, and 4.10, where skewness is represented, the
positioning of the means and medians has been somewhat altered to aid
visualization. If drawn to exact scale, the median would divide the area
under each curve into two halves of equal size. As for the mean, if the
curve were a solid object placed on a fulcrum, the object would be exactly
balanced at the mean, much like two people on a seesaw balancing one
another. In Figure 4.7, for instance, the actual locations of the mean and
median are farther to the left than they appear.
3. These figures originate from several sources and are summarized
in W. Spencer, Global Studies: The Middle East, 8th edition. Copyright ©
2000 by The McGraw-Hill Companies, Inc. All rights reserved. Reprinted by
permission of McGraw-Hill/Dushkin Publishing. [Link].
’ WA 1 osa
rye|

2 Ds sner S10 ss
W KEY CONCEPTS

mean deviation/average variance


measures of dispersion deviation/mean standard deviation
absolute deviation definitional formula
absolute value computational formula
CHAPTER 5

Measuring Dispersion

W PROLOGUE ¥
SN EEE PD SERENE STEN s Se a

Comparing two groups by a measure of central tendency may run the


risk for each group of failing to reveal valuable information. In particular,
information about the distribution of the scores within each group may be
useful to us but not revealed by the mean, median, or mode. In some
groups, the scores may all fall near the middle score, whereas in other
groups, the scores may be more widely spread above and below the central
scores. Accordingly, it is possible that the more bigoted group of the two we
compared, using a measure of central tendency, might contain some highly
bigoted individuals but possibly also several less bigoted people than could
be found in the less bigoted group. So in addition to central tendency, we
should examine the dispersion of the scores in each group as well.
POORER SIMI SISO BLS ASRS SEELEY
I LESS EEE
ESE TEE HIE SCTE ELAS TERNS TES EY AEE SGOT EE EDITED
MEME ESIESE
128 << STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
In addition to finding measures of central tendency for a set of scores, we
also calculate measures of dispersion to aid us in describing the data.
Measures of dispersion, also called measures of variability, address the
degree of clustering of the scores about the mean. Are most scores rela-
tively close to the mean, or are they scattered over a wider interval and thus
farther from the mean? The extent of clustering or spread of the scores
about the mean determines the amount of dispersion. In the instance
where all scores are exactly at the mean, there is no dispersion at all; dis-
persion increases from zero as the spread of scores widens about the mean.
In this chapter, we will cover four measures of dispersion: the range, the
mean deviation, the variance, and the standard deviation.

Measures of dispersion Measures of variability that address the degree of


clustering of the scores about the mean.

Dispersion The extent of clustering or spread of the scores about the mean.

VISUALIZING DISPERSION

To begin our discussion, let us suppose that in a penology class, three teach-
ing assistants—Tom, Dick, and Harriet—had their respective discussion
groups role-play court-employed social case workers who read the files of
convicted criminals and recommended to the judge the penalty to be imposed
for each criminal. The teaching assistants then compared each student’s rec-
ommended sentence to the one actually imposed by the real judge. The teach-
ing assistants then rated each student on a 0 to 10 scale, with 10 being a totally
accurate reproduction of the sentences that were actually handed down. There
were four students in each discussion group. The results were as follows:

Tom's Group Dick’s Group Harriet’s Group


Mi x= x=
: 9 10
2 8 10
38 7 6
> x =32 > x =32 ye =32

=e 32 a, 32
a, a 32 es
a 8 X Harriet — & = 8
eoEs 4 - ° Dick— 4
Measuring Dispersion » 129

The three groups share the same mean, but f=


the dispersion of the scores varies from none
in Tom’s group to some in Dick’s group to
even more in Harriet’s group. This is illus-
trated in the histograms to the left. Because
the distributions of individual scores clearly Oo
ES
Go
ho)
=
differed from each other in terms of their dis- Lee ned 54 uO canoe 9 10)
persion, we need to measure that dispersion Tom’s Group
in addition to measuring central tendency. fe
In this chapter, we will discuss measures
of dispersion in an order that will ultimately
bring us to the two measures used to the vir- INO
Gar

tual exclusion of the others, the variance and 1 |


its positive square root, the standard devia- 0
tion. The first two measures we will discuss, Pek oS ASS Oe era Ie10)
the range and the mean deviation, may be Dick’s Group
thought of as building blocks for understand-
ing the variance and standard deviation. Since
such measures are rarely used with data hav- }
ing a level of measurement less sophisticated 5
than interval level, they are usually calculated '
along with the calculation of the mean. With ; ‘
the mean as our measure of central tendency, 2s SSG 8) 9910
we then calculate a measure of dispersion,
Harriet’s Group
most often the standard deviation.

THE RANGE
The range is the simplest measure of dispersion. It compares the highest
score and the lowest score achieved for a given set of scores. The range can
be expressed in two ways: (a) with a statement such as “The scores ranged
from (the lowest score) to (the highest score),” or (b) with a single number
representing the difference between the highest and lowest score.

Range The simplest measure of dispersion that compares the highest score and the
lowest score achieved for a given set of scores.

In the case of Harriet’s group, whose scores were 6, 6, 10, and 10, we would
say, “The scores ranged from 6 to 10.” Or we could express the range as the dif-
ference between 6 and 10 (10 — 6) or 4. “The scores in Harriet’s group had a
mean of 8 and range of 4.” Now we can compare the ranges of the three groups.
130 << STATISTICS FOR THE SOCIAL SCIENCES

> Harriet’s Group: Scores ranged from 6 to 10. Range = 10-6 = 4.


> Dick’s Group: Scores ranged from 7 to 9. Range = 9-7 = 2.
» Tom’s Group: Scores ranged from 8 to 8. Range = 8- 8 = 0.

These ranges correspond to the spread on the histograms for the three
groups, with Harriet’s group’s scores being most dispersed about the mean,
Dick’s being less dispersed, and Tom’s having no dispersion at all.
Although we commonly make use of the range in our day-to-day
discourse, it really is not a very meaningful measure of dispersion. Because
only the highest and lowest scores are taken into consideration in finding
the range, the other scores have no impact. Just as in the case of the mean,
where an extreme value of x can distort the mean and lessen its usefulness,
the use of only the extreme values can render the range less useful. Our next
measure, the mean deviation, rectifies this situation.

THE MEAN DEVIATION


The mean deviation (V.D.) (also called the average deviation or the
mean absolute deviation) is sensitive to every score in the set. It is based
on a strategy of first finding out how far each score deviated from the mean
of the scores (the distance from each score to the mean), summing these
distances to find the total amount of deviation from the mean in the entire
set of scores, and dividing by the number of scores in the set. The result is
a mean, or “average,” distance that a score deviates from the mean.

Mean deviation An average distance that a score deviates from the mean.

To get the mean deviation, we first find the distance between each score
and the mean by subtracting the mean from each score. Let us use Harriet’s
group as an example.
Harriet’s Group
x= X= x-X=
10 8 fs
10 8 2
6 8 —2
6 8 —2
Measuring Dispersion > 131

At this juncture, we encounter a problem: We cannot add up the x-—x


column to get the total amount of deviation in the system. Recalling that the
mean is the value ofx that satisfies the expression Y>(@&—X) = 0, we can see
that if x = 8, adding algebraically, the x — Xs for each student in Harriet’s
group produce a sum of zero:

SiGe a9)
= 242-2, 2 4 A=
This is because the positive deviations (where x is greater than the mean)
exactly balance the negative deviations (where x is less than the mean).
Recall that we currently are seeking the distance from each score to the
mean, without regard to direction; that is, we do not care whether x is
greater or less than x. Like a car’s odometer, we want to count the distances
traveled, disregarding the direction or directions in which we drove. We do
this by taking the absolute value of each x — x, the distance disregarding
its sign (in effect treating all x —X s as if they were positive numbers). We
symbolize the absolute value of a deviation as |x —x|. When we add up all
these absolute values, $~ |x — X|, we get the total amount of deviation
of the scores from the mean. When we divide that sum by the total number
of scores, we get the “average” amount (the mean amount) that a score
deviated from the mean of all of the scores: the mean deviation.

Absolute value The distance or difference disregarding its sign. Here, the distance
between each value of x and the mean, regardless of whether x is greater than the
mean (a positive distance) or less than the mean (a negative distance).

Thus,

wee
|x —x|
n

For Harriet’s Group:

C= x xX-X= xX —x| =

10 8 2 2
10 8 2 Z
6 8 —2 2
n=4 8 —2 2
yon 182 wile asa 8

x—Xx 8
eee ig ripe bes ear?)
132 << STATISTICS FOR THE SOCIAL SCIENCES

For Dick’s Group:

Xx X-xX = jn-X|=
1 1
0 0
0 0
ca dl
oe see

32 x —X 2
mp. = X& | =
ae
oe — 8
n 4

For Tom’s Group:

X= x—-x = lxn-—x|=
8 0 0
8 0 0
8 0 0
n=A4 BS ee)
eei
osm
(eo) 0 0
yx= 32 Sea 2

ey ak 0
mp. = =! ies
ie
x= — = 8

These results are in keeping with our expectations: Harriet’s group has
the largest mean deviation, Dick’s has a smaller one, and Tom’s has the
smallest (a value of zero).

THE VARIANCE AND STANDARD DEVIATION

The formula for the variance resembles that of the mean deviation except
that S~ |x —X| is replaced by the expression )* (x —X)’. Instead of taking
the absolute value of each deviation, we square it to get rid of negative
numbers. (Remember that a negative number times itself is a positive
number, just as a positive number times itself is a positive number.) Since
the squares of the deviations greater than one unit will be much larger
than their respective absolute values, 5° (x —xX)* will usually be larger than
)~ |x -X|, and the final variance will usually be larger than the mean devi-
ation. To adjust for this and produce a result more comparable to the
Measuring Dispersion ® 133

mean deviation (more like an “average” amount of deviation), we often


take the positive square root of the variance, thus producing the standard
deviation, indicated for now by the letter s.

Thus,

.
x —XxWw
Variance = s* = 2 =x)
n

IY (6 — x)
Standard Deviation = s = LG =x)?
n

Variance An “average” or mean value of the squared deviations of the scores


from the mean.

Standard deviation The positive square root of the variance, which provides
a measure of dispersion closer in size to the mean deviation.

Let us calculate s* and s for our three groups—Tom’s, Dick’s, and


Harriet’s—whose mean deviations were 0, 0.5, and 2.0, respectively.

Jom’s Group

es x= LEX = (x-x)?=
8 8 0 0
8 8 0 0
8 8 0 0
8 8 0 2)
Ge) 0

Thus,

ee ORE OP er
n 4

The variance and standard deviation both equal zero, as does the mean devi-
ation, for this group in which there is 70 dispersion at all.
134 << STATISTICS FOR THE SOCIAL SCIENCES

Dick’s Group

ve x= Mak = (ia ie
9 8 1 1
8 8 O O
8 8 0 0
yf, 8 —| il
ore iy =2

Thus,

eS Cie)
oe eee
Regsh I (es
n oe:

pay x—-xe [2 =)fl 2onr20n


22 2/2 207
n 4 2

Remember that it is the standard deviation (0.7), not the variance, which
substitutes for the mean deviation (0.5).

Harriet’s Group

Oo x = Ce-X = C=)
10 8 Z 4
10 8 2 4
6 8 —2 4
6 8 2 4
> @=x)* = 16

Thus,

= =—=40
n 4

Yi @ —x)? /16 -
5= =4/— = V4 = 2.0
n 1

Let us compare our measures. See the histograms at the top of the next
page.
Measuring Dispersion je 135

f= Range = 0
4 Mean Deviation = 0
3 Variance = 0
Standard Deviation= 0
1

0
(eee
Ge 7 aonaG
Tom’s Group
f= Range = 2.0
4 Mean Deviation = 0.5
3 Variance = 0.5
; Standard Deviation = 0.7

1
0 :
i BE ale 5 oe Hehe 12s AIC)

Dick’s Group

f= Range = 4.0
Mean Deviation = 2.0
Variance = 4.0
Standard Deviation = 2.0

oo
Oo
Ne
IP 2 394556)
7% 8 TO

Harriet’s Group

Below are the dispersion measures for artistic freedom for the non—
liberal arts majors, Group A, presented in Chapter 4.

GroupA
x= x= x-xX= |xn-x| = (x =x)?=
8 7 1 1 1
8 v7 1 1 1
8 7h 1 il 1
w/, 7 O 0 0
7 vi 0 0 0
7 x 0 0 0
6 7 —l 1 ]
6 V —] 1 1
n=9 6 c = i 4
Silerser ate y. @&-x)*= ON
x= 63°

The scores range from 6 to 8. Range = 8 - 6 =2.


136 STATISTICS FOR THE SOCIAL SCIENCES

Pee 2 ey,
n 2

i ee ee
9
Beane
Variance
= s* = es = : = 4 = 0.67
n 9 e)

Standard Deviation = s = V0.67 = 0.82

Summary Group A
Range 2.00
Mean Deviation 0.67
Variance 0.67
Standard Deviation 0.82

As mentioned, the variance and standard deviation are the most widely
used measures of dispersion in statistics, even though on the face of it, the
mean deviation would appear to be the most logical measure (and easiest to
calculate) of the three. The reason is that the standard deviation has mean-
ing in terms of a common frequency distribution known as the normal
curve, which we will encounter later in this text.

THE COMPUTATIONAL FORMULAS


FOR VARIANCE AND STANDARD DEVIATION

The variance formula s* = > & —x)*/n is often referred to as the defini-
tional formula since it not only calculates the variance but also defines or
explains what the variance is: the mean amount of the squared deviations of
the scores from the mean. (It is often quite difficult for those long away from
algebraic formulas to “see” that definition, but it is there.)

EAN

Definitional formula A formula that not only calculates the variance but also
defines or explains the concept. In the case of the variance, the formula defines it as
the average (mean) amount of the squared deviations of the scores from the mean.
erases nant ssoame

For computational purposes, however, it is often easier to use one of sev-


eral alternative formulas, known as computational formulas, particularly
if a calculator is available. One such computational formula is the following:
Measuring Dispersion 187,

Computational formulas A formula that generates a correct answer but does not
seek to define what the concept, such as the variance, actually is.

a\2
Six? — ie
Variance = s* =

Standard Deviation = s =

Before we apply these formulas, we should make note of the difference


between two parts of the formula: )°x* and (}°x)’, which are vot the same.
The first, }°x*, read “summation of xsquared,” tells us to square each x and
then add up all of the x’s. The second, (})x)’, read “summation of «x,
quantity squared,” tells us to first add up all the xs to get)’x and then
square ) \x to get ()
\x)’. (This follows the convention of first doing what is
inside a set of parentheses before doing what is outside of the parentheses.)
Thus, we must add the original scores and square the sum, and we must also
square each original score and add up the squared values.

Group A

x= x= ‘
8 64 rie — aN 447 — OS"
8 64 iam 7 =a 9
8 64
y 49 baits aks ae _ 447-441
7 49 a 9 a 9
7 49 2
6 36 = Q ee Gy
6 36 <TR
n=9 6 36 and

SS SoCo ia Be eS 2) s= 4067 = 0.82

Open 63)
= 65°65
= 3969
The answers are obviously the same as when we use the definitional
formula. Often, the two results will differ slightly due to rounding error,
particularly if the mean used in the definitional formulas is not a whole
number (such as 7, in this case) but possesses several decimals (such as 7.2,
138 << STATISTICS FOR THE SOCIAL SCIENCES

BOX 5.1

Another Formula for the Standard Deviation

In Chapter 8, you will encounter another formula for the standard


deviation, indicated by the lowercase Greek letter sigma with a circum-
flex above it and read (believe it or not) as “sigma hat.”

Note that this formula is the same as the definitional formula we have
just been using except that 7 — 1 replaces 7 in the denominator. When
we wish to generalize about some group (called a population) from
data taken from fewer people than the entire group (called a sample),
we run into a problem. Suppose I wanted to generalize about the ages
of all residents of Thousand Oaks, California (the population), from a
sample of 20 residents of that town. If 1 calculate the mean for my sam-
ple, I get the best estimate of the mean age of all that community’s res-
idents that my data will allow. However, if I estimate the population’s
standard deviation from my sample, using the formula with 72 in the
denominator, my estimate is inaccurate. In fact, the smaller the size of
my sample, the less accurate my estimate of the population’s standard
deviation will be.
It turns out that the formula with 7 — 1 in the denominator gives
us a better estimate of the population’s standard deviation than the
formula with 7. Thus, you will see the 72 — 1 formula widely used in text-
books, calculators, and computer programs. In fact, rarely can we study
whole populations directly; so much of the time, we are really using
sample data to estimate population data. That is why the formula with
nm — 1 in the denominator appears so often.
Finally, note that many authors will state that the formula with 7 in
the denominator is for a population’s standard deviation and the 7 — 1
formula is for a sample’s standard deviation. That is not quite correct,
but since most of the time what we really are doing is using sample data
to estimate population data, we really are not interested in the sample’s
standard deviation except as an estimate of the population’s stan-
dard deviation. So, it is easier just to call the 7 — 1 formula the formula
for a sample’s standard deviation. That practice is not followed in this
textbook.
Measuring Dispersion BY)

7.23, 7.234, and so on). Notice that the computational formula requires the
calculation of several large intermediate figures, such as the (S°x)° = 3909.
Since such large numbers are not needed when using the definitional for-
mula, we may question the need for a computational formula. If, however,
there are many scores (even as few as the 9 scores in Group A), it is faster
and easier to use the computational formulas. It is even easier to use the
computational formulas with today’s advanced scientific, business, and
statistical calculators, which usually store ix and ees in their memories
for easy retrieval.

VARIANCE AND STANDARD DEVIATION


FOR DATA IN FREQUENCY DISTRIBUTIONS

If the data are in frequency distributions, the formulas given above will not
find the correct variance or standard deviation. In a frequency distribution,
we must account not only for each possible value of x but also for the
number of times, or frequency, that value occurs. This is the same reason we
modified the formula for finding the mean ofa frequency distribution in the
previous chapter. Recall that in calculating the mean for the liberal arts
majors, Group B, we first established an fx column and added it up to get
y_ fx. We then divided )° fx by }*f(our 7) to get the mean. For frequency
distribution data, the definitional formula for the variance is also adjusted
so that before adding the squared deviations, we multiply each squared
deviation by the frequency of that particular value of x.

Pie ae
ign n yi

Therefore,

Group B
an fo= | X= x-K= (&-X)= (x -x)f=
oe 2 Ses ays 1.5 see 225 = 50
8 5) 24 | 7.5 0.5 O25 O255635 = 0075
7 5) Shel eS) =05 025 ~O0.25*3 = 075
6 2 i =15 DOS en =

i ae 10 yf = 7/5 SlGgeer ml == 00)


140 << STATISTICS FOR THE SOCIAL SCIENCES

Thus, the variance is

MEDICS
a ee, Due
eae ee
on eeeel
icsn es
n Dor 10

and the standard deviation is

s =V1.05 = 1.0246 = 1.03


For data in frequency distributions. there is also an adjusted computa-
tional formula.

extf= Qupor2 2: mea


i =a
f=

n De
To apply this to Group B, we must generate columns for x? in order to find
yox’ and xin order to find )°x?f. We have already generated an fx column,
but we need to square its summation.

ie dis pos | x = xf
9 Zi 18 | 81 81x 2S 162
8 3 24 | 64 64x 3=192
z 6 21 | 49 49 x 3 = 147
6 Re 12 | 36 36K 2S 72
n= f=10 Sfe=75 ara 5
Qh fey’ = (75)?
=75x75
= 5625
Thus, the variance is

oe ef — we _ 52 UO) ag7a_ 3625


P i ee 10
7. 10 10
573 — 562.5 10.5
10 10

and the standard deviation is

6 ony AO

The results are identical to those found using the definitional formulas,
Measuring Dispersion » 141

We now know the primary measures for describing a single-interval or


ratio-level variable: the mean for central tendency and the standard devia-
tion or variance for dispersion. With the latter two, we generally use the
standard deviation for descriptive purposes but retain the variance for use
in procedures that will be discussed later in this text.
With the exception of the range, the measures of dispersion presented
in this chapter all assume interval level of measurement. (The range may
be applied also to ordinal data: “The guests at the $100-a-plate charity
fund-raiser ranged from middle class to affluent.”) While measures of dis-
persion are widely used with interval-level data, they are only rarely used
with lower levels of measurement. Accordingly, such usage will not be
covered here.

CONCLUSION
We have now covered the last of the basic tools of descriptive data analysis.
With the introduction of dispersion measures, particularly the variance and
the standard deviation, we can begin the study of several statistical tech-
niques widely applied in many disciplines. We will see that in addition to
their role as useful descriptive tools, the mean and the variance often plug
into other formulas. Thus, they do double duty. Armed with the tools intro-
duced so far, we will eventually return to the task of finding and describing
relationships between two variables.

Chapter 5: Summary of Major Formulas

Individual Data

The Mean Deviation

pesIX
— Xx
n
|
The Variance The Variance
Definitional Computational

ye Yo (x — Xx)? i yx? — 7"


NN nl
142 < STATISTICS FOR THE SOCIAL SCIENCES

Frequency Distributions

Definitional Computational

x2 oe (ofc)?
Py i oe ee ees,
- n Py 8 5 n Dake
Both Individual and Frequency Distribution
The Standard Deviation

S = V the variance

Bata@ ts:
Note: For the following exercises, refer to the exercises at the end of Chapter 4 for
the definitions of the variables.

Exercise 5.1
In the social worker sample (Exercises 4.7 to 4.9), a group of 9 private agency
employees was compared to a group of 16 public employees. Following are the
health care cost ratings for the private agency employees. Remember that the
higher rating indicates more concern about the issue.

Private Agency Employees


Health

Find the mean Health score.


Find the median.
Find the mean deviation.
Find the variance using the definitional formula.
Find the variance using the computational formula.
NH
WwW
Bb
GF
NO
= Find the standard deviation.
Measuring Dispersion ® 143

Exercise 5.2
Following are the health care cost ratings for the public employees:

Public Employees
Health

95
95
95
90
90
90
90
90
90
80
80
75
7)
60
40
2D

Form a frequency distribution from the above, and using the appropriate formulas:

1. Find the mean Health score.


Find the median.
Find the variance using the definitional formula.
Find the variance using the computational formula.
Find the standard deviation.

UF
OS
bwCompare the mean and standard deviation of the public employees to those of
the private agency employees found in Exercise 5.1. Which group’s scores clus-
ter more closely about its mean?

Exercise 5.3
Management personnel have been scored on a scale measuring assertiveness
of leadership style, where more assertiveness indicates less accommodativeness.
Are financial and banking managers more assertive than their colleagues in other
service industries? Following are scores for 7 managers in finance- or banking-
related firms.
144 STATISTICS FOR THE SOCIAL SCIENCES

Assertiveness

24
49
92
92
ul
68
7

Find the mean Assertiveness score.


Find the median. (Note that you must first array the data from high to low scores.)
Find the mean deviation.
_ Find the variance using the definitional formula.
Find the variance using the computational formula.
Po
Co)Find the standard deviation.

Exercise 5.4
Following are assertiveness scores for 18 managers from nonfinancial service
industries listed in an ungrouped frequency distribution.
x = Assertiveness f=
100
oF
OF
86
54
30
27
24
22

um
Ow RS
Oye
SB
aS
Ww
Nt
SS
|S
NO]

Find the mean Assertiveness score.


Find the median.
Find the variance using the definitional formula.
Find the variance using the computational formula.
Find the standard deviation.
wNn
DOR
Compare the means and standard deviations of the nonfinancial institution
managers to those found in Exercise 5.3. Which group is more assertive? Which
group’s scores are more spread out about the mean?
Measuring Dispersion » 145

Exercise 5.5
Below are the results, in printout format, for the employee sample of Exercise 4.10
(refer to Exercise 4.10 for a definition of the variables). Please note that this was run
using SAS, one of several statistical packages available (we will be discussing the
most recent version of SAS later in this book). Like most such packages, data are
presented with far more decimal places than social scientists need. While suitable
for engineers and some scientists, this level of precision is not suitable for the less
exact measures that we use. Thus, when discussing the results, we will round to
one or two decimal places.
In this exercise, workers have been broken down by region, Midwest versus all
other regions combined. Suppose it had been rumored that the corporation was
planning to close several plants and move those jobs to plants in other countries
with lower wage scales. Suppose it had also been rumored that only plants in the
Midwest would be exempt; in all other regions, some plants would be shut down.
Let us compare the attitudes of the employees.

Reg = Midwest
Variable N Mean SD.
ATTEND 13 90.6153846 122782902
BOARD 13 44.7692308 19.2663517
DIV 13 76.6153846 16.8302231
SECUR 1 67.7692308 28.4580213
PARTIC 13 39.6153846 35.0868885
OPPOR 13 55.4615385 38.5413531
UNION 13 55.3846154 35.5844968
SALARY 13 65.6923077 25.9466909

Reg # Midwest
Variable N Mean 6.0.
ATTEND 87 93.7297297 5.8720082
BOARD BF 34.7837838 18.1615209
DIV 37 78.7837838 16.1832134
SECUR a7 44 702702/ 32.6129361
PARTIC 37 67.4324324 30.5646315
OPPOR 37 30.0270270 32.4349661
UNION 37 76.8918919 29.4380553
SALARY 37 49.6486486 27.4764710

1. Compare the means for each variable. What do you conclude?


2. Which region usually has the greater diversity on these dimensions as
determined by comparin g the standard deviations ? In which two scales is
that tendency reversed?
146 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 5.6
Following is a comparison of the managerial group to the employee group.
MGTPOP
Variable N Mean SD.
ATTEND 89 923595506 9.9307228
BOARD 89 971123596 15.1840513
DIV 89 74.9438202 16.4305422
SECUR 89 56.1685393 32.4479429
PARTIC 89 48.5955056 34.6737126
OPPOR 89 42.5280899 36.3509065
UNION 89 62.4719101 31.6307136
SALARY 89 53.8764045 24.2575294

EMPLOY
Variable N Mean SD.
ATTEND 50 92.9200000 8.0997899
BOARD 50 37.3800000 18.7832861
DIV 50 78.2200000 16.2081989
SECUR 50 50.7000000 32.9274093
PARTIC 50 60.2000000 33.7602592
OPPOR 50 36.6400000 35.5486214
UNION 50 71.3000000 32.2118307
SALARY 50 53.8200000 27 F501 67

You have already compared the means in Exercise 4.10.


Now compare the standard deviations for each variable. What can you conclude?
For which variables are the managers more diverse (have larger standard deviations)?
For which variables are the employees more diverse?

Exercise 5.7
The two discontented groups, upper-middle management and white-collar
employees, are compared in the following sets of data.

UPPER-MIDDLE MANAGEMENT
Variable N Mean S.2:
ATTEND 50 91.8000000 12.8364914
BOARD 50 472200000 10.9195388
DIV 50 78.6400000 15.1600442
SECUR 50 39.8000000 31.0227040
PARTIC 50 72.1000000 22.0178499
OPPOR 50 16.4800000 18.2043278
UNION 50 85.6600000 10.8130873
SALARY 50 36.1000000 11.1158097
Measuring Dispersion > 147

WHITE-COLLAR EMPLOYEES
Variable N Mean SD.
ATTEND 29 93.3103448 SS OE
BOARD 29 23.8965517 6.9710873
DIV 29 84.2068965 9.4354169
SECUR US 31.3103448 26.7956598
PARTIC 29 81.8965517 17.8992115
OPPOR 29 1321724137 16.7333477
UNION 29 92, 4482758 10.9628796
SALARY 20 35.0000000 14.7672417

Compare the means and then the standard deviations for each variable. What do
you conclude?

Exercise 5.8
For the data in Exercise 4.1, calculate and compare the standard deviations. Use
_ the definitional formula to find the variance for the exporters and the computa-
_ tional formula to find the variance for the nonexporters. Then find and compare the
two standard deviations.

Exercise 5.9
For the data in Exercise 4.4, calculate and compare the standard deviations. Use
the frequency distribution definitional formula to find the variance for the exporters
and the frequency distribution computational formula to find the variance for the
nonexporters. Then find and compare the two standard deviations.
WV KEY CONCEPTS V
ea SAE AAS EIN BE OI Rem NSE iS penntagseces PAL NS

contingency table spurious relationships antecedent variable


control variable causal models intervening variable
ea WD MERE DEOL LE ELIE LEER RENE NIE ILL EER LEI EER SSG
CHAP TILER 6

Constructing
and Interpreting
Contingency Tables

VY PROLOGUE V¥

In chugter 1, we reduced es notion a contingency ables or


cross-tabulations (or just tables). Remember the example comparing sky
conditions to the presence of rainfall. In this chapter, we return to tables
and specifically how they are used with social science data. Recall that a
table is designed to test a hypothesized relationship between two grouped
variables, usually nominal or ordinal level of measurement. Specifically,
we will use as examples the exploration of possible relationships between
economic development (modernization) on one hand and _ political
development (democratization, civil and/or political rights) on the other.
Are these two concepts related? Is democracy associated with economic
development?
PERE OLENA EEO LALO LLL IEEE IDOLEIIL,

Pe 149
150 @ STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION

In this chapter, we further develop the subject of tables. We begin by


discussing how to further recode the variables in a data set to create useful
tables that relate two of the variables from that data set. We will also present
procedures for changing the table’s entries from frequencies to percent-
ages, which makes most tables easier to interpret. We will look at examples
of both the construction and the interpretation of cross-tabulations.
In the latter part of the chapter, we will discuss the impact that a third
variable might have on the relationship between two other variables, in
particular, the impact of the third variable on tables relating to the first two.
We will also discuss how to imply a direction of causation among the vari-
ables and will look at the variety of causal relationships that may be found
within a set of three variables.

CONTINGENCY TABLES
In Chapters 1 and 2, we introduced the logic behind the interpretation of
tables. Contingency tables (or cross-tabulations or just tables) continue to
be tools for both interpretation and presentation of data. Consequently, con-
tingency tables are often presented in scholarly research even if far more
complex statistical techniques of analysis were also used by the researcher.
Thus, table construction and interpretation remain essential components of
data analysis.
The contingency table depicts a relationship between the inde-
pendent variable and the dependent variable. Each variable is in grouped
data format and is generally nominal or ordinal level of measurement, even
if the original data were individual interval level (as explained in Chapter 2).
If this is to be the case, once the variables have been selected, the researcher
returns to the original data and creates grouped categories, often with
new labels such as high, medium, or low. In effect, these newly created cat-
egories reflect a new operational definition on the part of the researcher
since the researcher determines the cutoff points for high, medium, and
low. Since tables with nominal and ordinal variables were presented earlier,
we will now emphasize such ordinal variables resulting from grouping
individual interval data.

Pe ATONE
SI PE ASHRAM EERE FEDER SOURED
USHA ONE SOMERS

Contingency table Table that depicts a possible relationship between the


independent variable and the dependent variable.
me tent atteCCRT SETEASE NAN ASE
DRE SEIS ANSTO
Constructing and Interpreting Contingency Tables » 151

Here, we have selected a sample of 20 countries and have selected data


about them from two major sources. The political data come from Freedom
in the World 2003 ([Link]), and the economic and social
data come from The Economist: Pocket World in Figures, 2004 edition
(London: Profile Books, Ltd., 2003). The Corruption Perceptions Index was
developed by Transparency International ([Link]) and was
taken directly from its Web site. Let us examine the relationship between
two variables: Per Capita Gross Domestic Product (GDP/Capita) and Political
Rights. Both appear in the data list in Table 6.1. In doing so, be advised that
a small convenience sample of 20 countries was selected so that you could
easily envision the process of table construction from original data. Because
this sample of 20 countries may not be representative of all the countries of
the world, consider any conclusions from this data set to be only suggestive
of what one might find in studying all countries. The issue of sampling will
be taken up again later in this text.
We wish to test the hypothesis that the level of wealth enjoyed by a
nation’s citizens is positively related to that nation’s level of political rights.
We know that historically, those countries with competitive democratic
political systems and high levels of political and social rights and freedoms
have been among the world’s wealthier industrialized countries, though
that is not exclusively true. To see how much of a relationship exists, we use
per capita GDP as the independent variable.
The dependent variable is an index of political rights originally
developed by Raymond Gastil. Gastil’s index originally ranged from 1 to 7,
with the score of 1 going to the countries with the highest levels of political
rights. To simplify this example, countries are grouped into a three-category
ordinal scale with categories of High, Medium, or Low (H, M, and LI).
The independent variable, GDP/Capita, is the total value of all final goods
and services produced within a nation’s economy divided by the size of the
population—a widely applied index of relative wealth.

REGROUPING VARIABLES

Although the political rights index has already been grouped, we will need
to group per capita GDP There are two basic ways to approach this prob-
lem. One is to see if there is already some external standard in use. For
instance, the International Monetary Fund or the United Nations may have
already created operational definitions for high, medium, and low per capita
GDP If this is the case, we use those definitions.
If no such precedents exist, the way we group the data will be based solely
on our own preferences and priorities. Although there are no hard-and-fast
uoynndog
auogdaja,
OCSpe SEY81 i OL PLSHE LYS ey S67EV yil6 BLS

O01 ad
EL

SAULT
|e

= 40fddD
uoywonpy
suypuads
LY LY og xe LY eG 8g ly cE Py) OS DY SS, SL Sav

JO % SB
suipueds

S38 a) 68 SPS 9°6 sie S'6 6S 0'8 8's Is) Sie Ges EL
40fdd)
be
JO % SB

qyveHE
aanynaasy —Uvgal)%
dd5Jo%

sumiy p> uoyvindog

Gz v8 ee vy LL VOLic CVC V1 v8 cio GE DY OC 60


WOM

OL
SAOJDIIPUT JVIIOS PUD JYI11JOd JO YOOGPUDH PJAOM WO

716Lass:682Le SSLGY,SSL6 682 6'$8S69)OrGieLES£8 S68


LZ
douvjoadxg

0'C8GE 618ge CL ONE8'C89°F9 [$8 £08082Tel LOSOs £08


QVUa]
afi]

ti)
O8h
J'6L
£6S
O'S
VY ME
v6
es
OLY
O18
816 SY
VVC
Lc
£8
79
Om OT
oF
$99
66L0'9¢
DLL
van eSre
ee
FO
6T
9°7E

ddD
vudvo

OL0°6I

06¢ CZ

O¢0'7Z
OLF

Ore1
OLSH

007 SE
OOL
Aad
006

09
O¢6L
OCF
O16z

OFLZ
06SZ
OSS‘LI
068°97

ZE
OTS‘

OSL*EZ
0767
suoydarsad
uUOo’NdnALoyD

COOE XEPUT

pane
9°8

ae}
9€

ERG
Co

S6

Sy
OY

£6
ae
ci

iY
0'6
0'F

LEG
ol
PE

He

a
AN
AN

AN
dd

dd
d

d
d

‘gary Ayfensed = yg ‘d0JJ 10U = AN :92°J = J “ALON


dd

WLOPaaL]
Q'S
09
C9

OF

09
OT

ai
CT
OT

OT
oi
OF
aOT

dOt

JOT
SG

dSe

dOT
dOe

AOl
d

£00E
SULIDY

[2119
SVT ve Pensed

Sn IN rete SENSIS te teal hea GNY ost: axe—t GN Lig tos AP Am ve \O

Sasa]
[0211

mr lie FOR NTR AHH AAA


0d
SIG51N
saqris
91GBU9

voy yInos
rIquinjoy

pouuy,
UdPIMS
epeuey)

a0uely

puelod
eISsny
1dABy

uedef
[IZe1g
L

purjaly
eIpU] [OBIS] ial IMGUQUINZ
Pury)
eYesny

eAUdy
MON
pur[eaz

152 <<
Constructing and Interpreting Contingency Tables >» 153

rules for doing this, the following procedure is helpful. Look at the range
of scores (lowest to highest) and create class intervals along this range. Do a
frequency distribution of scores and examine that distribution. Then combine
some of the class intervals with two objectives in mind: (a) reducing the
number of class intervals to a manageable number (three or four categories)
and (b) having in each category at least a few cases and, if possible, a roughly
equal number of cases in each class interval. Rarely will the final breakdown
meet the criteria for grouped interval-level data as discussed in Chapter 2.
In this case, we end up with grouped ordinal data.
To ease our task, we change the scale to reflect units of $1000. (That is,
we divide each GDP by 1000. Thus, for example, the U.S. score of $35,200
becomes a score of 35.2.) We do not have to do this, but data presentation is
easier if we do. Now, we note that the scores range from a low of .360 (Kenya)
to a high of 35.2 (United States). A preliminary frequency distribution might
yield five categories, labeled—arbitrarily—from very high to very low:

GDP/Capita
(in thousands
of dollars) Range f= ps

Very High 30-35.2 || 2


High 20-29.9 leh 5
Medium 10-19.9 ||| 3
Low 1-9.9 HLL 6
Very Low Below 1 Bae 4

We may reduce the number of ranked categories, although in this case,


we cannot do so and keep our category sizes roughly equal. Let us combine
very high and high:

GDP/Capita
(in thousands
of dollars) Range f=

High 20-35.2 7
Medium 10-19.9 5)
Low 1-9.9 6
Very Low Below 1 4

Throughout this process, the researcher chose class interval sizes, cut-
alterna-
off points, and category names. You might easily have chosen other
tives to the ones selected here.
154 @ STATISTICS FOR THE SOCIAL SCIENCES

We used a similar procedure for the Political Rights index to produce a


four-category scale.

Political Rights Rank f=

High 1 12
Medium 2-3 Ze
Low 4-5 a,
Very Low 6-7 )

We are now ready to construct our table. To do so, we follow a series


of conventions used for table construction. (You will find in the research
literature many tables that deviate from some of these conventions. What
follows is the way tables are generally constructed.) By convention, we gen-
erally place the independent variable so that its categories become the
columns of the table, and the categories of the dependent variable become
the table’s rows (see Table 6.2).

Table 6.2

GDP/Capita

High Medium Low Very Low


Political Rights 20.0-35.2 10.0-19.9 1.0-9.9 Below 1.0

High (1) HTT || ||


Medium (2-3) | |
Low (4-5) || |
Very Low (6-7) | ||

Once we have laid out the table, we examine each country in Table 6.1,
go to the column of the table within which that country’s GDP/capita falls,
and go down the column to the row within which that country’s Political
Rights score can be found. In the cell where the appropriate column and
appropriate row intersect, we put a tally mark. We continue with the process
until all 20 countries have been tallied.
For example, from our original data, we see that the United States has
GDP/capita of $35,200, which, after dividing by 1000, gives a score of 35.2,
which falls in the High category (20.0-35.2). We also see that the United States
ranks High for Political Rights. We find the cell where the High GDP column
intersects the High Political Rights row and place a tally mark in that cell.
When we have finished all 20 countries, we will have a total of seven tally
marks in that (High-High) cell, corresponding to Canada, France, Ireland,
Japan, Sweden, the United Kingdom, and the United States. In the lowest
Constructing and Interpreting Contingency Tables » 155

right-hand cell (Low GDP/Capita and Very Low Political Rights), we will have
two marks, corresponding to China and Zimbabwe.
As we discussed in Chapter 1, a clear clustering appears along the main
diagonal of the table, indicating the presence of a positive relationship
between the variables.
We replace the tally marks in each cell with the sum of the tally marks and
put in marginal totals and the grand total for reference, as shown in Table 6.3.

Table 6.3

GDP/Capita

High Medium Low Very Low


Political Rights 20.0-35.2 10.0-19.9 1.0-9.9 Below 1.0 Total

High (1) y 3 2 0 12
Medium (2-3) 0 0 1 1 Z
Low (4-5) 0 0 2 1 3
Very Low (6-7) 0 0 1 2 3

Total e 3 6 4 20

GENERATING PERCENTAGES

When the category marginal totals of the independent variable are very
close to one another, it is possible to interpret a table from the cell entries
alone, as we did in Chapter 1. Often, however, this is not the case. The cate-
gory totals are often larger numbers, and they differ, sometimes considerably,
from one another. Accordingly, it is almost always advantageous to change
the table from frequencies to percentages before interpreting the data.
Percentaging in this manner has the effect of creating an equal number
of cases in each category of the independent variable, so that trends in the
table may be more readily observed. It is as if there were 100 high GDP/
capita countries, 100 medium, and 100 low. With equal numbers of cases,
comparisons across the categories are easier to make.
Consider the following table with frequencies as its cell entries. Both the
top and bottom rows increase as one scans from left to right.

If the table is percentaged so that each column adds up to 100%, however,


the relationship is far more apparent.
156 @ STATISTICS FOR THE SOCIAL SCIENCES

80% 15%
20% 85%
100% 100%

Percent means “per 100,” and the idea behind percentaging is to make
the individual cell entries add up to a common base, a total of 100%. There
are three ways to base our percentages (and most computer programs do
all three for us, whether we need them or not).

1. Treat each cell entry as a percentage of the grand total. For instance,
we have 7 cell entries in the High-High cell of Table 6.3 and a grand total
of 20.

7 2035
= 35%

If we percentaged similarly for each of the 16 cells in the body of the


table and then added all of these percentages together, they would add up
to 100%. In the analysis of tables, however, this kind of percentaging is rarely
useful. We really are not interested in the fact that 35% of the countries in
our study had high GDPs/capita and also high levels of political rights.

2. Treat each cell entry as a percentage of its row total. In the High-
High cell, 7 out of a total of 12 countries have high political rights levels.

7 + 12 = 5833 = 58.33%

Thus, 58.33% of the countries with high political rights scores also have
high per capita GDPs. We repeat the process for the other GDP categories
in the same row and add up the three percentages for the row.
Countries with a high political rights score having:

High GDPs/Capita 58.33%


Medium GDPs/Capita 25.00%
Low GDPs/Capita 16.67%
Very Low GDPs/Capita 0%

Total of all countries with a 100.0%


high level of political rights

We then do the same for the other rows.

3. Treat each cell entry as a percentage of its column total. There are a
total of seven countries with high per capita GDPs, all of which have high
political rights levels.
Constructing and Interpreting Contingency Tables » (a7

7 + 7=1.000 = 100.0%

No country is in the medium, low, or very low political rights category.


Adding these up, we get
Countries with high (20.0 and above) GDPs per capita having:

High Political Rights Scores 100.0%


Medium Political Rights Scores 0
Low Political Rights Scores 0
Very Low Political Rights Scores 0

Total of all countries with high 100.0%


GDPs per capita

The convention generally followed is to percentage so that each


category of the independent variable sums to 100%. The independent
variable is usually the one whose categories are the columns of the table,
and thus percentaging is usually such that each column of the table sums
to 100%.
Once our percentaging is done, we will be able to compare across
categories of per capita GDP Without percentaging, differences among the
scores in each GDP group could merely reflect the different overall sizes of
these groups. With percentaging, all GDP categories have the same overall
size, namely, 100. Thus, if within a single category of the dependent variable
the percentage in each GDP group differs, we know that the differences are
likely due to a relationship between the variables.
We need one more thing to complete our table—a title. There are many
possible formats for the title, but the format often encountered in the liter-
ature will be used here. First, name the dependent variable, followed by a
comma and the word by, and then name the independent variable. Other
supporting information, such as year and location of the study, may also
be inserted parenthetically where appropriate. In simple form, the title of
Table 6.4 would be as follows:

Political Rights, by Per Capita GDP

We could expand this as follows:

Political Rights Levels, by Per Capita GDP (in thousands of U.S.


dollars), c. 2001, for Selected Countries (in percentages)

The title for Table 6.4 is a compromise between the two.


158 < STATISTICS FOR THE SOCIAL SCIENCES

Table 6.4 Political Rights, by Per Capita GDP (in thousands of dollars)
(in percentages)

GDP/Capita

High Medium Low Very Low


Political Rights 20.0 and Above 10.0-19.9 1.0-9.9 Below 1.0

High 100.0% 100.0 33.3 0


Medium 0 0 16.6 25.0
Low 0 0 a 25.0
Very Low 0 0 16.6 50.0

Total 100.0% 100.0% 99.8% 100.0%


(72 =) Y) (3) (6) (4)
SOURCE: Freedom in the World 2003 (New York: Freedom House, [Link]);
The Economist: Pocket World in Figures, 2004 edition (London: Profile Books, 2004).

Note four other things about Table 6.4. First, a minor point: It is only
necessary to put one percentage sign (%) in the body of the table, in the
upper left-hand cell. (Some people include the percentage sign in each of
the entries of the top row to indicate that percentages add to 100% for each
column.) Also, include percentage signs by the 100.0% totals at the bottom.
Second, note that the percentage total for the low GDP group is only
99.8%. This is simply due to rounding error. We will often be a bit below or
a bit above 100% due to rounding error, but be careful. If it is adding to 96%
or 104%, then Houston, “we have a problem.” There’s more than just
rounding error here. Third, a major point: We always report the number of
cases (7 =) for each category of the independent variable. This is because
most tests and measures that we will learn to calculate from tables will be
based on numbers of cases, not percentages. By reporting the 7s, we can
always retrieve the data found in Table 6.3, the original cell entries. For
instance, in the High Political Rights, Low GDP cell, we find 33.3% of the
Low GDP countries have High Political Rights. At the bottom of that col-
umn, in parentheses, we see that there are six Low GDP countries. Thus,
33.3% of 6= .333 x 6 = 1.998 (due to rounding error) = 2 = the cell entry in
Table 6.3.
The final point to note about Table 6.4—one of great importance—
is the footnote citing the sources of the data and (where possible) page
references. Here, a reference is given to a bibliography at the end of the
article where the complete citation would be provided. Note: Failure to cite
sources is plagiarism.
Constructing and Interpreting Contingency Tables » 9)

INTERPRETING

In the following discussion, we will review some of the points presented in


Chapter 1. A positive relationship will be indicated by a clustering on the main
diagonal (upper left to lower right). An inverse relationship will be indicated
by clustering on the off diagonal (lower left to upper right). Now, however,
we have percentages in our table instead of the original frequencies.
We interpret our table by examining increases or decreases in per-
centages of the dependent variable as we move from one category of the
independent variable to the other. In the case of a 4 x 4 table set up so that
the upper left-hand cell represents the largest quantity on both variables
(High-High) and the lower right-hand cell is the least quantity for both
variables (Very Low-Very Low), if the relationship is positive, we expect
approximately the following:

Peed GDP/Capita
Rights High Medium Low Very Low

High Highest % 2nd Highest % 3rd Highest % Lowest %


in this row — in this row in this row in this row
Medium
Low eae eee ee ae
Very Low Lowest % 3rd Highest % 2nd Highest % Highest %
in this row — in this row in this row in this row

The pattern in the Medium Rights category is harder to interpret.


Generally, we would expect less change across this row than across any
other. If there had been more than three categories of Political Rights, the
changes in intermediate categories may be vague. For instance, if a category
existed between High and Medium (e.g., “Medium-High Political Rights”),
changes across that row could be similar to those in the High row but less
severe, or the changes could be quite small, as in the Medium row.
This sort of pattern in a table is not the only pattern that indicates a rela-
tionship, but it is the most commonly found pattern and is called a linear
relationship. We will discuss linear relationships again later.
Beginning at the top row (High Rights) of Table 6.4, we look at the per-
centages from left to right: 100% to 100% to 33.3% to 0%. Clearly, this goes
from the highest to the second highest to the lowest percentage. In the
Medium Rights row, the percentages (from left to right) read as follows: 0%
to 0% to 16.6% to 25.0%. For the Low Rights row, we expect a turnaround,
follow
with the percentages increasing from left to right, and that row does
160 @ STATISTICS FOR THE SOCIAL SCIENCES

the expected pattern to some extent, 0%, 0%, 33.3%, but then down to
25.0%. But for Very Low Political Rights, the pattern is a clearly increasing
one: 0%, 0%, 16.6%, 50.0%. With a few inconsistencies, therefore, Table 6.4
suggests a positive relationship between the variables.
If the relationship were imverse, we would expect the upper row to
increase from left to right and the lower row to decrease as shown below:

High Lowest % 3rd Highest % 2nd Highest % Highest %


Medium ———__—— ————— a
Low a
Very Low Highest % 2ndHighest% 3rd Highest % Lowest %

The percentages in such a table might look like this:

High Medium Low Very Low


High 276 42.8 65.5 705
Medium 1S 15.2 13.3 15.0
Low 10.0 5.0 ee 2
Se
Very Low 77.0 37.0 Dove 11.0
100.0% 100.0% 100.0% 100.0%

Suppose there is ”o relationship between the variables. Then, for each


row, we would expect the percentages to be the same (or nearly the same).
The following is an example of no relationship:

High Medium Low Very Low

High 60.0% 60.0 60.0 60.0


Medium 20.0 20.0 20.0 20.0
Low 5.0 5.0 5.0 5.0
Very Low 15.0 15.0 15.0 15.0

Here, 60% of all countries have high levels of political rights regardless of
GDP/capita; 20% have medium political rights, again regardless of GDP; 5%
have low rights; and 15% have very low rights for all four categories of GDP

Example

The 11 variables presented in Table 6.1 are presented below.


Operational definitions for these variables have been arbitrarily established
by assigning scores to three or four selected class intervals for each example.
A recoded data list for the sample of 20 countries (Table 6.5) accompanies
the operational definitions.
uoyvjndog
OOL 42d
Gea Getta tesSstas
S Sat r za

SUIT
40f dd)
{0 % Sv

=
auogdajay
supuads
40f dd)

SaAeaHrtAteawATtsss al
uUOoYvINpY
G1jVaH
SB
%{0
sulpuads

ddd
Jo
uoyvindog

% WOM
QANYIN
IASPUY
SULUIY
UDG)
douvpedxq

Loe See aS Se ee ae TA
% afi]
(ssuryues

vydvo
aVUulay]
dd)D
suoydarsaqd
UOUdNALOD
Aad

SHOTS AA ITS TSsrssAvAArtIrrsS


padnoss)[euIpIO

Dies eee ee eh cn eh ee
suljvy

Xopuy
c00E

Bess eS fe SS Sista Sa eS
SAMAAGIT

£00E
ee ast]

WOpaa]
papoddy

Heo Ss ie iS SS Se res Sra aS


i) jP2uyOd
SIGSIM

nS SS Snes aoe Se eS
IMGVQUNTZ

pur[eaz
21981$°9

epeury

MON
OUR]

pouuy
euly)

idAsyq

saivig
SO)

[Zerg
elyelasny [avIS]
purely
PIpU] vAUdy purjog
uedef eIssny
yInos
eoLyy
UIPIMS ‘USIH
‘WNIpo|]
VAALON
HWJTA
‘MOT
AIDA
=
rIquinjo?)

161
162 << STATISTICS FOR THE SOCIAL SCIENCES

Per Capita GDP (in $1000 U.S.) i=


High 20.0—35.2 7
Medium 10.0-19.9 3
Low 1.0-9.9 6
Very Low Below 1.0 4

Political Rights j=
High 1 12
Medium 2-3 2
Low 4-5 >
Very Low 6-7 >

Civil Liberties v=
High 1 8
Medium 2-3 6
Low 4—5 3
Very Low 6-7 5

Freedom Rating f=
High 120 8
Medium 1.5-3.0 6
Low ean! ©,
Very Low 5.5 and above 5

Note: Freedom House uses a different categorization: Free (1.0—2.5),


Partly Free (3.0-5.5), and Free (5.5—7.0). Based on additional criteria, some
countries with scores of 5.5 are rated partly free and others are rated not
free. These are indicated in Table 6.1.

Corruption Perceptions Index f=


High (perceived 7.59.5
least corrupt)
Medium 6.0-7.4 4
Low 4.0-5.9 S
Very Low (perceived 1.9-3.9 z
most corrupt)

Female Life Expectancy j=


High 80.0-85.1 8
Medium 70.0-79.9 8
Low 45.6-69.9 4
Constructing and Interpreting Contingency Tables » 163

Percentage of Population That Is Urban f=


High 80.0-91.8% 6
Medium 70.0-79.0 6
Low 50.0-69.0 3
Very Low 32.6-49.9 .
Percentage of GDF: Based on
Agriculture and Mining =
High 15.0-25.1% 5
Medium 4.0-14.9 i
Low 1.4-3.9 8
Health Spending as Percentage of GDP f=
High 9.0-13.0% 4
Medium 7.0-8.9 9
Low 3,.8-6.9 ca
Education Spending as Percentage of GDP f=
High 6.0-10.4% 5
Medium 5.0-5.9 4
Low 4.0-4.9 8
Very Low 2.1-3.9 3

Telephone Lines Per 100 Population f=


High 50-73.9 dh
Medium 20-49.9 6
Low 0-19.9 el

Experts in economic and political development argue that the percent-


age of the GDP based on agriculture in a country’s economy is a good index
of industrialization, of modernization, and, thus, of wealth. The idea is that
the less economically developed a nation, the greater its agricultural work-
force. To test this, assume Percentage of GDP Based on Agriculture to be the
independent variable and Per Capita GDP to be the dependent variable,
since it is an estimate of the nation’s wealth. We set up a table as a tally sheet
and fill it in with information from the data list of Table 6.5. On the data list,
the 20 countries are listed alphabetically down the left-hand side, and the
variables are listed along the top. Australia, for instance, is assigned to the
medium category for per capita GDP and the medium category for percent-
age of GDP based upon agriculture (and mining). In Tables 6.6 and 6.7,
Australia would be one of the two countries in the medium-medium cell.
The other country in that category, surprising or not, is New Zealand.
164 @ STATISTICS FOR THE SOCIAL SCIENCES

Table 6.6 Percentage of GDP Based on Agriculture (and Mining in the Case of
Australia)

High Medium Low

Per Capita GDP 15% and Above 4.0-14.9% 1.4-3.9%

High (20.0-35.2) | HITT


Medium = _(10.0-19.9) ||
Low (1.0-9.9) | I |
Very Low (below 1.0) I |

Table 6.7 Percentage of GDP Based on Agriculture

High Medium Low

Per Capita GDP 15% and Above 4.0-14.9% 1.4-3.9%

High (20.0-35.2) 0 1 6
Medium (10.0-19.9) 0 2 1
Low (1.0-9.9) 1 4 1
Very Low (below 1.0) 4 0 0

Total > H 8

A clustering appears in the off diagonal, indicating the anticipated


inverse relationship: As percentage in agriculture increases, GDP/capita
decreases. In Table 6.7, we replace the tallies with the actual frequencies.
From the frequencies in Table 6.7, we calculate the percentages shown
in the final table, Table 6.8. The percentage changes, particularly in the high
and low Per Capita GDP rows, indicate the existence of a relatively strong
inverse relationship between the two variables.

CONTROLLING FOR A THIRD VARIABLE


We have been dealing so far with the relationship between only two vari-
ables at a time. At this point, we explore the influence of a third variable,
called the control variable, on the relationship between the first two
variables. Later, we will discuss the exact role played by that third variable;
that is, is it also an independent variable or is it another dependent vari-
able? For now, however, let us merely examine some of the influences of
a third variable.
Constructing and Interpreting Contingency Tables » 165

Table 6.8 Per Capita Gross Domestic Product for Selected Nations, by
Percentage of GDP Based on Agriculture

Percentage of GDP Based on Agriculture

High Medium Low

Per Capita GDP 15% and Above 4,0-14.9% 143.9%

High (Z010=35.2) 0% 14.3 75.0


Medium _(10.0-—19.9) 0 28.6 abs:
Low Gio-9)) 20.0 5V1 12:5
Very Low (below 1.0) 80.0 0 0

Total 100.0% 100.0% 100.0%


(n=) (6) (7) (8)

Control variable A third variable that may have an influence on the relationship
between the first two variables.

What we are doing is making use of statistical analysis in a manner anal-


ogous to a scientific experiment in a laboratory setting. In many experiments,
one group of subjects, known as the control group, is given no experimen-
tal treatment but is compared to others who do receive treatment. If we are
studying the effects of time constraints on decision making, one experimen-
tal group may have 10 minutes in which to make a decision. A second exper-
imental group may have 20 minutes to do so. But a third group, the control
group, would have no time constraints placed on it. If all three groups reach
similar decisions, the researcher could argue that time constraints have no
influence on decision making. If one or both experimental groups reach dif-
ferent decisions than the control group, the researcher could argue that time
constraints do have an influence.
At the conclusion of this chapter, Exercise 6.2 is a table in which there is
a marked inverse relationship between the percentage of the labor force
engaged in agriculture and the number of telephones per 1000 people. Why
is this the case? Probably since telephones are more likely to be found in
wealthier industrialized societies, which, by their nature, have fewer people
employed in agriculture than do preindustrial societies. We might look for a
control variable such as per capita GNP (wealth) or proportion of people in
metropolitan areas (urbanization) and explore what happens to the original
relationship when the effects of the control variable are taken into account.
If the original relationship turns out to be the result of the presence of wealth
166 STATISTICS FOR THE SOCIAL SCIENCES

or income, we say that the original relationship between agriculture and


telephones is an indirect relationship and, specifically in this case, a spurious
relationship, caused by the presence of the control variable. We use the
term spurious when the relationship between two variables is the product of
a common independent variable. Thus, agriculture and telephones are func-
tions of a single independent variable—wealth. If in time order, wealth did
not come first but followed agriculture and preceded telephones, we would
say wealth provides interpretation of the relationship between agriculture
and telephones. By the same token, suppose that among low per capita GNP
countries, the same relationship between agriculture and telephones exists,
as can be found among high per capita GNP countries. In this case, wealth
would appear to have vo influence on the original relationship.

Spurious relationship — Relationship between two variables that is the product of a


common independent variable.

Let us use an example in which the control variable functions as a


second independent variable. We want to look at what happens to the
relationship between the dependent and independent variable when we
“control for” the influence of this new variable. By control for, we mean
to explore the relationship between the initial two variables for each of the
categories of our new control variable. In doing so, we can see the impact
of the control variable on the initial relationship.
Suppose we conduct a study for a sample of 200 newly commissioned
military officers. One of the variables we examine has to do with the extent
that these subjects favor armed intervention abroad under circumstances
where the country’s vital interests are in jeopardy. We think that the gender
of the respondents will influence their attitudes on intervention. The
following table is generated:

Altitude on —
Military Intervention Male Female
Favor 70% 60%
Oppose 30% 40%
Total 100% 100%
(=) (100) (100)

There appears to be some relationship between gender and attitude,


with the males a bit more inclined toward favoring intervention. We now
wish to see to what extent, if any, this relationship is changed with the
Constructing and Interpreting Contingency Tables j» 167

introduction of a control variable. Since our study is being done in the


United States, we choose race (white, nonwhite) as the control variable.
In other countries, different variables might affect the relationship. For
example, if these were Canadian Armed Forces personnel, we might use
language (Anglophone, Francophone) as the control variable. In the United
Kingdom, we might use region (England, Scotland, Ulster, Wales). We use
race in this example in order to work with an easy-to-represent dichotomy.

PARTIAL TABLES
There are two general outcomes that may ensue when the control variable
is entered.

1. The control variable has no impact on the initial relationship.

2. The presence of the control variable changes the initial relationship


or is necessary for there to be a relationship between the indepen-
dent and dependent variables.

We look first at the case where the control variable has no impact. To illustrate
this, we generate two partial tables. Each one shows the relationship
between attitude and gender for a specific category of the control variable.
Assuming 120 white and 80 nonwhite respondents, we will have three tables:
the initial table, the table for white respondents, and the table for nonwhite
respondents. The first set of tables will be in actual frequencies (see Table 6.9).

Table 6.9

Race

Gender White Nonwhite


Attitude on ————
Intervention Male Female Male Female Male Female

Favor 70 60 42 36 28 24
Oppose 30 40 18 24 2 16

Total (7 =) 100 100 60 60 40 40

Note that if you add the corresponding frequencies in the partial tables
together, you retrieve the original data. Thus, adding the 42 white males
who favor intervention to the 28 similarly inclined nonwhite males yields 70,
the frequency in the initial table on the left.
168 STATISTICS FOR THE SOCIAL SCIENCES

To interpret the tables, we now generate the percentages, as shown in


Table 6.10. The percentage is the same for each corresponding cell: 70% of
all males favor intervention, as do 70% of white males and 70% of nonwhite
males. In this example, race has no effect on the original relationship.

Table 6.10

Race

Gender White Nonwhite


Attitude on
Intervention Male Female Male Female Male Female

Favor 70% 60 70 60 70 60
Oppose 50) 40 30 40 30 40

Total 100% 100% 100% 100% 100% 100%


(n =) (100) (100) (66) (60) (40) (40)

When the presence of a control variable changes the initial relationship,


there are several possible outcomes, creating partial tables from the same
initial table that we have been using. First is a case where the partial table
percentages differ from the original ones, but the strength of the relation-
ship remains about the same within each category of the control variable
(see Table 6.11).
Second, we have a case where the relationship is stronger in one cate-
gory of the control variable (here, among whites) than in the other category
(see Table 6.12).
Third, we have a case where the relationship exists for one category of
the control variable but not for the other. Here, the relationship disappears
for nonwhites (see Table 6.13).
Finally, a very strange pattern emerges in Table 6.14.
In Table 6.14, the original relationship all but disappears when we control
for race. To understand this phenomenon, examine the frequencies shown in
Table 6.15, from which percentages were generated in the partial tables.
Note the small number of whites, regardless of gender, who Oppose
intervention. The original relationship is not really between gender and atti-
tude but between race and attitude. Whites overwhelmingly favor interven-
tion, while nonwhites overall tend to oppose intervention. To see more
clearly that race, not gender, is what causes attitude to vary, let us recon-
struct the table showing the relationship between race and attitude, We add
the 54 white males who favor intervention to the 36 white females who favor
it, giving a total of 90 whites in favor. We repeat this for those who Oppose
and get the frequencies shown in Table 6.16.
Constructing and Interpreting Contingency Tables >»

Table 6.11

Race

Gender White Nonwhite


Attitude on
Intervention Male Female Male Female Male Female

Favor 70% 60 80 70 60 50
Oppose 30 40 20 30 40 50

Total 100% 100% 100% 100% 100% 100%


(n =) (100) (100) (60) (60) (40) (40)

Table 6.12

Race

Gender White Nonwhite


Attitude on
Intervention Male Female Male Female Male Female

Favor 70% 60 80 67 >» 50


Oppose 30 40 20 33 45 50

Total 100% 100% 100% 100% 100% 100%


(n =) (100) (100) (60) (60) (40) (40)

Table 6.13

Race

Gender White Nonwhite


Attitude on
Intervention Male Female Male Female Male Female

Favor 70% 60 83 67 50 50
Oppose 30 40 17 33 50 50

Total 100% 100% 100% 100% 100% 100%


@=) (100) (100) (60) (60) (40) (40)
170 STATISTICS FOR THE SOCIAL SCIENCES

Table 6.14

kace

Gender White Nonwhite


Attitude on = ——
Intervention Male Female Male Female Male Female

Favor 70% 60 93 92 38 39
Oppose 30 40 fi 8 62 61

Total 100% 100% 100% 100% 100% 100%


(n=) (100) (100) (58) io?) (42) (61)

Table 6.15

Race

White Nonwhite
Attitude on
Intervention Male Female Male Female

Favor 54 36 16 24
Oppose 4 3 26 37

Total 58 a) 42 61

Table 6.16

Race
Attitude on
Intervention White Nonwhite

Favor 90 40
Oppose 7 63

Total 97 103

Percentaging
Favor 93% 39
Oppose Gi real

Total 100% 100%


(n =) 97) (103)
Constructing and Interpreting Contingency Tables » 171

Table 6.17

Race

Gender White Nonwhite


Attitude on ee
Intervention Male Female Male Female Male Female

Favor 50% 50 100 0 0 100


Oppose 50 50 0 100 100 0

Total 100% 100% 100% 100% 100% 100%


Ci) (100) (100) (50) (50) (50) (50)

In this situation, where the original relationship between attitude and


gender was based on race, we call the original association indirect: The rela-
tionship between attitude and gender was really due to the racial factor—
not gender at all. We use partial tables such as those above, as well as partial
correlation coefficients (see Chapter 14), to help us discover spurious and
other relationships.
Another possible outcome of a control variable may be a case whereby
a relationship between two variables may not be apparent except in the pres-
ence of the control variable. This is illustrated in Table 6.17. Note that unlike in
previous examples, the initial table before controlling shows no relationship.
The original lack of relationship was brought about by two offsetting
relationships. Among whites, all males favor and all females oppose inter-
vention. Among nonwhites, all females support intervention and all males
oppose it. This corresponds to the one instance mentioned in Chapter 1
where association would not be necessary for causation. The offsetting rela-
tionships must balance one another, but they need not be perfect relation-
ships (see Table 6.18).

Table 6.18

kace

Gender White Nonwhite


Altitude on
Male Female Male Female Male Female
Intervention

50% 50 80 20 20 80
Favor
50 50 20 80 80 20
Oppose
100% 100% 100% 100% 100%
Total 100%
(n=) ~ ~€60) (100) (50) (50) (50) (50)
172 @ STATISTICS FOR THE SOCIAL SCIENCES

There are other occurrences that we might encounter with partial tables.
It is possible under certain circumstances that there is no original association,
but upon partialing, each partial table shows a small relationship. Unlike the
above relationships, however, the partial tables’ relationships do not offset
one another but actually run in the same direction. Also, we may occasionally
encounter an initial table with some amount of association in it and, upon
partialing, find that the association in both partial tables runs in the direction
opposite the association in the initial table. These two cases are rare enough
that we need not illustrate them here.

CAUSAL MODELS
The findings in the partial tables are combined with other assumptions to
produce causal models of these relationships. Often, we portray these
models by use of schematic diagrams indicating the independent, depen-
dent, and control variables—/, D, and C, respectively. Lines are drawn
between each pair of variables having association. (Of course, if there is no
initial association, no line is drawn.) If the association is later proven to be
indirect, it is replaced by a dotted line. Finally, an arrowhead is placed on the
line to indicate the probable direction of causality. Since D is the variable
being explained, at least one arrow must point to that variable. Thus, our
diagrams are based on observation and logic. We observe the relationship
between each pair of variables to determine whether a line should be
drawn between them and whether the line should be solid or dotted. Logic
determines the direction of each arrowhead: If D is dependent, what is the
logical flow of causation?

Causal models Schematic diagrams showing the independent, dependent, and


control variables and, where appropriate, positing the flow of causation of change in
the dependent variable.

Figure 6.1
In the example we have been using, D is attitude on military
intervention, J is gender, and C is race. Since gender and race
develop at about the same time biologically and attitudes are
shaped by both factors roughly concurrently, it is logical to assume
D that both gender and race are independent variables acting on atti-
a tude. Accordingly, a schematic of the relationships in Tables 6.11,
6.12, and 6.13 could look like the one in Figure 6.1.
Constructing and Interpreting Contingency Tables » 1)2433

For Table 6.10, where race had no impact, the schematic might Figure 6.2
look like the one in Figure 6.2.
For Table 6.14, where the /—D relationship was indirect, we |
could use the schematic in Figure 6.3.
A double-headed arrow between two variables (or two parallel |
arrows pointing in opposite directions) could indicate reciprocal — |
causation between two variables (see Figure 6.4). However, be {| ~~
aware that certain techniques using such models exclude the ‘es
option of reciprocal causation and require the selection of a single
direction for each arrow. Figure 6.3
One last point: In our example, race (C) and gender (/) develop
concurrently; what if that were not the case? Suppose C were not I
race but socialization, the process whereby attitudes (including |
those pertaining to armed intervention) are learned and internal-
ized. Our model then might be similar to the one in Figure 6.5. |

Figure 6.4 C

Hostility 9 <¢@————___________» | Aggression


Figure 6.5

Gender “determines” the type of values to which one is social- |


ized, and socialization shapes attitudes toward a specific military
intervention. (Assumption: Boys are led to believe that it is best to
fight back; girls learn that peaceful solutions are preferable.) Here,
we say that gender, /, is antecedent, and socialization, C, is inter-
vening. That is, / leads to D by way of C. C
EEE
NOC ERSTE ONO

Antecedent variable The variable initially leading to change in the dependent


variable.
Intervening variable The variable through which the antecedent variable brings
about change in the dependent variable.
NILES LESS Svea

Assuming no double arrows, Figure 6.6 demonstrates several possible


models for a three-variable situation with D being dependent. However, this
only scratches the surface of the problem. Advanced techniques enable us
to explore situations where there are several control variables working
simultaneously and the number of possible models increases rapidly. For
our purposes, though, looking at just one control variable is sufficient to
appreciate the variety of ways external factors can influence a relationship
between any two variables.
174. STATISTICS FOR THE SOCIAL SCIENCES

Figure 6.6

| and C are concurrent:

C is antecedent to |:

C is intervening:

C is subsequent to D:

C is irrelevant to the | - D relationship:

| is antecedent to both C and D

C is both antecedent to and concurrent with |:

| is both antecedent to and concurrent with C:

| is both antecedent to and concurrent with D:


Constructing and Interpreting Contingency Tables >» Wo

COMPUTER APPLICATIONS
In this text, three sets of library programs for generating computer-driven
output will be used. Two will be introduced here, SPSS 12.0 and SAS 9.1.
Later we will use Microsoft Excel. All three are designed for use using
Microsoft Windows. All three are menu driven and relatively easy to oper-
ate. These are only variations of programs available from these companies
and from many other firms. The examples presented here are illustrative
and widely used in the social sciences, but by no means are they exhaustive.
Your instructor may be modifying the information presented here for some
other set of computer programs.

SPSS A set of statistical computer routines: Statistical Package for the Social Sciences.
SAS_ A set of statistical computer routines and a programming language:
Statistical Analysis System.
Microsoft Excel Microsoft's spreadsheet program that also may be used for
Statistical analysis.
Microsoft Windows — Microsoft's widely used operating system.

Because of space considerations, the examples run in this book will be


simple ones, only touching some of the capabilities of these systems. You
are encouraged to consult the manuals available for each system as well as
other texts available in your computer center or campus library.
As I said above, my goal is to get you to the statistical output you
need as quickly as possible, skipping whatever frills or refinements may
be skipped for now. What always amazes me is the ease by which accurate
output is attained today on a personal computer, without even needing
anymore to use a mainframe or facing the mechanical problems of earlier
eras. (One of these days, I'll tell you young whippersnappers about the time
that the counter-sorter chewed up my entire data deck and left it as a pile
of confetti on the computer center’s floor.)

SPSS

An opening screen on SPSS gives you several options. Of these, click on


Type in Data and then click ok.
You have the option of typing in your data first, in which case the vari-
ables will be numbered as below: VARO0001, VAROOO02, etc. As an option you
may use your own variable names. At the bottom left of the screen, you will
see data view and variable view. Click on variable view. Click on the box to
the right of the 1 under name and type in GDP and to the right of number 2,
under GDP type in Agriculture. Spaces available and use of symbols are
176 @ STATISTICS FOR THE SOCIAL SCIENCES

limited. For now, the examples used under Computer Applications will retain
the variable numbers.
VARO0001 will be GDP/Capita, and VARO0002 will be Percentage of GDP
From Agriculture. Recode the information from Table 6.5 into numerical
equivalents, as follows:

VAROOOOL VAROOO02

H 1 A 1
M Z M 2
L ) Ly 3
VL 4

The numerical codes replace High, Medium, Low, and Very Low, respec-
tively. These codings appear as ordinal rankings in order that the computer-
generated table will resemble Tables 6.7 and 6.8. (After you have done this
run, try recoding VAROO001 as H4, M3, L2, and VL1 and see what happens to
the table generated by the program.) Once done, your data list should
appear as it does in Table 6.19.
If you examine Table 6.19, you will notice that case number one
(Australia) is coded 2.0 for both VAROO001 (GDP/Capita) and VAROOO02

Table 6.19

VAROOOOL VAROOO02

2.00 2.00
3.00 2.00
1.00 3.00
4.00 1.00
3.00 2.00
3.00 1.00
1.00 3.00
4.00 1.00
1.00 2.00
OANIAWKRWND 2.00
3 OM 3.00
|— 1.00 3.00
4.00 1.00
a 2.00 2.00
Od
DN
~
3.00 3.00
aoA 3.00 2,00
16 3.00 2.00
1.00 3.00
=a
NJ
co 1.00 3.00
19 1.00 3.00
i)oO 4.00 1.00
Constructing and Interpreting Contingency Tables >» WHA

(Percentage of GDP From Agriculture and Mining). This corresponds to the


Medium scores of Australia for both variables in Table 6.5.
Click Analyze at the top of the screen and each subsequent item from
the lists that will appear thereafter. For the whole procedure, click

Analyze
Descriptive Statistics
Crosstabs

A Dialog box will appear on the screen. In a smaller box on the left,
VAROO0001 and VAROO002 will be indicated under it. To the right of that box
are two buttons with > symbols in them. You'll click on each button to move
VAROOOO01 to the Rows box to the right of the top button and VAROO002
to the Columns box to the right of the bottom button. Thus, VAROO001,
GDP/Capita, becomes the row variable, the dependent variable, and
VAROO002, Percentage of GDP From Agriculture and (in the case of Australia)
Mining, becomes the column variable, the independent variable.
There are several other buttons at the bottom of the Dialog box. One of
these, statistics, we will leave alone for now. (In a later chapter, we will redo
this run, adding many additional features to our table, but for now we keep
things simple.) Now, click the ce//s button, and after that, click column per-
centages where indicated. We are telling the program that, in addition to the
frequencies for each cell in the table, we want the percentages to add to
100% for each column in the table. By doing this, we will duplicate the infor-
mation found in Tables 6.7 and 6.8.
Return to the Dialog box by clicking continue. Leave the format button
alone for now. Click ok to run the program. The table will appear on your
computer screen. It should be the same as Table 6.20.

SAS

Once you have opened the SAS program, look along the top of your
screen for the word solutions and click on it. A list will appear to the right.
Click on the first entry of that list, avalysis, and another list will appear to its
right. Click the second entry in that list, a@7alyst. To summarize, you have typed

solutions
analysis
analyst

Once you have done this, a data list appears. Under A on that list, you
will enter the GDP/Capita data, and under B on the list, you will enter the
Percentage of GDP From Agriculture and Mining data. When done, your
screen should mirror Table 6.21.
178 @ STATISTICS FOR THE SOCIAL SCIENCES

Table 6.20 Crosstabs

Cases

Valid Missing Total

N % N % N %

VAROOOO1 * VAROOOO2 20 100.0% 0 0% 20 100.0%

VAROO001 * VAROOOO2 Cross-tabulation

VAROOO002

1.00 2.00 3.00 Total

VAROOOO1 1.00 Count 0 1 ef


% within VAROOOO2 0% 14.3% 75.0% 35.0%

2.00 Count 0 2 a
% within VAROOO002 0% 28.6% 12.5% 15.0%

3,00 Count i 4 6
% within VAROOOO2 20.0% 57.1% 12.5% 30.0%

4.00 Count 4 0 4
% within VAROO002 80.0% 0% 0% 20.0%

Total Count 5 i 20
% within VAROOOO2 100.0% 100.0% 100.0% 100.0%

At the top of the screen, click on statistics, and a new list will appear to
the right. Click the second entry on the list, table analysis. To summarize,
you now have typed

statistics
table analysis

On the left-hand side of your screen, over the word remove is a box with
the letterA in the top line of the box and the letterBunder it. Highlight (left
click) the A. Then, in the center of the screen, note the word row. Click on
that word, and A will disappear from the remove box and appear in the box
below the word row. You have now indicated thatA, GDP/Capita, is the row
variable, the dependent variable.
To the right of the word row is another button containing (you guessed
it!) the word column. Go back to the remove box and highlight B. Then click
Constructing and Interpreting Contingency Tables » Wy

Table 6.21

A B G 1D) E
1 2 2
2 3 2
3 1 3
4 4 1
5 3} 2
6 3 1
i 1 3
8 4 1
9 1 2
10 2 3
11 1 3
12 4 1
13 2 2
14 3 3
15 3 2
16 3 2
dy 1 3
18 1 3
19 1 3
20 4 1

on the word column, and B will appear in the box under the word column
[Link] from the remove box. Percentage of GDP From Agriculture
and Mining is now the column variable, the independent variable.
In case this seems redundant, bear in mind that in most studies,
you have more than two variables, so in the data list, unlike the one in
Table 6.21, you would also have data in columns C, D, and so on. For
each job, therefore, you must identify the dependent and independent
variables and, if appropriate, whatever control variable(s) you are using.
(With SAS, you indicate any control variables by clicking them into the
strata box.)
You are actually ready to click the ok button to submit your run, as
column percentages are a default option in SAS, so you don’t have to spec-
ify them. (To verify this fact, click the tables button and note that column
percentages are already checkmarked. Click the ok button to return to the
previous screen.) Back at the screen where you indicated the row and
column variables previously, click ok. Your table will appear on the screen.
It should be identical to Table 6.22.
180 << STATISTICS FOR THE SOCIAL SCIENCES

Table 6.22 The FREQ Procedure Table of A by B

A B

Frequency
Col Pct 1 a 5 Total

1 0 1 6 yi
0.00 14.29 75.00

Z 0 Zs 1 3
0.00 28.57 12.50

) 1 4 1 6
20.00 57.14 12.50

4 4 0 0 4
80.00 0.00 0.00
Total 5 7 8 20

CONCLUSION

Tables traditionally have been handy tools for presenting findings and
demonstrating relationships between variables. They are relatively easy
to learn to interpret, and data at any level of measurement may be put into
tabular form. They still represent a common form of data presentation,
although perhaps less so than in the past.
As we have seen, partial tables enable us to investigate the impact of
a third variable on the relationship between two other variables. Moreover,
partial tables provide a very thorough way of studying the impact of a
control variable. However, partial correlations—which will be covered in
Chapter 14—and other techniques are also used for similar purposes today.
The major weakness of tables is that unless there is either a perfect
relationship or a total lack of a relationship between variables, tables are
vague. As we have seen, we can spot a less-than-perfect relationship in a
table, but we cannot specify the extent of that relationship. In Chapter 11,
that weakness will be addressed, and the concept known as a measure of
association will be introduced. The measure of association provides a
number that seeks to reflect the actual degree of association of the variables
in a table.
Constructing and Interpreting Contingency Tables » 181

EXERCISES
Exercise 6.1
Formulate one or two hypotheses from the variables in the data list from Table 6.5
(other than those used below), create your own tables, and interpret the results.
~

Exercise 6.2
Two percentage tables are presented below. Write a short paragraph interpreting
each table.

Table E6.2.1 Political Rights Scores by Corruption Perceptions Index 2002

Political Rights Corruption Perceptions Index

High Medium Low Very Low


Scores Ve ee ee 6.0-7.4 4.0-5.9 EO43.9

High 1 100% 100 66.67 0


Medium 2~3 0 0 23.95 14.29
Low 4—5 0 0 0 42.86
Very Low © 6-7 0 0 0 42.86

Total 100.00% 100.00% 100.00% 100.00%


(7 =) (6) (4) (3) (7)
NOTE: The higher the Corruption Perceptions Index, the less corruption is perceived
to exist.

Table E6.2.2 Telephone Lines Per 100 Population by Percentage of GDP


From Agriculture
ES
Nc ee 27 ro RS ZI

Percentage of GDP From Agriculture

Telephone Lines High Medium Low


Per 100 Population 15.0-25.1% 4,0-14.9% — 1.4-3.9%
Se Se a a Re a
High 50-73.9 0% 143 75.0
Medium 20-49.9 0 57.1 25.0
Low 0-19.9 100.0 28.6 0

Total 100.0% 100.0% 100.0%


(n =) (5) (7) (8)
OO
182 << STATISTICS FOR THE SOCIAL SCIENCES

Table E6.3 Table of GDP/Capita by Percentage of GDP From Agriculture

Crosstabs

Case Processing Summary


Cases

Valid Missing Total

N % N % N %o

VAROOO001 * VAROOO0O2 20 100.0% 0 20 100.0%

VAR00001 * VARO00002 Cross-tabulation

VAROOOO2

2,00

VAROO001 1.00 Count


% within VAROOOO1
% within VAROOO02
% of total

2.00 Count
% within VAROOOO1
% within VAROOOO2
% of total

3.00 Count
% within VAROOOO1
% within VAROOOO2
% of total

4.00 Count
% within VAROOOO1 0% 100.0%
% within VAROOOO2 0% 20.0%
% of total 0% 20.0%

Total Count 7 20
% within VAROOOO1 25.0% 35.0% 40.0% 100.0%
% within VAROOOO2 100.0% 100.0% 100.0% 100.0%
% of total 25.0% 35.0% 40.0% 100.0%
Constructing and Interpreting Contingency Tables > 183

Exercise 6.3
Tables 6.7 and 6.8 could have been produced by a computer, such as the
accompanying table from SPSS. Variable 00001, the dependent (row) variable,
is GDP/Capita, and variable 00002, the independent (column) variable, is
Percentage of GDP Based on Agriculture and Mining. The codes are (1) high,
(2) medium, (3) low, and, in the case of GDP/Capita, (4) very low. In each cell, you
will see the frequency and, under it, three percentages based on the row total,
the column total, and the grand total, respectively.
Suppose percentage of the labor force engaged in agriculture was the
dependent variable and per capita GNP the independent variable. How would
you interpret this table?

Table E6.4 Table of Civil Liberties Score by Political Rights Score

Case Processing Summary

Cases

Valid Missing Total

N % N % N %o

VARO0004 *VAROOO03 =.20 1000% 0 0% 20, 1000%

VAR00001 * VARO0002 Cross-tabulation

VAROO003

1.00 2.00 3.00 4.00 Total

VAROO004 1.00 Count 8 0 0 0 8


% within VAROOO03 66.7% 0% 0% 0% 40.0%

2.00 Count 4 2 0 0 6
% within VAROOOO3 33.3% 100.0% .0% .0% 30.0%
3.00 Count 0 0 3 0 3
% within VAROOOO3 0% 0% 100.0% 0% 15.0%

4.00 Count 0 0 0 3 3
% within VARCOO03 0% 0% 0% 100.0% 15.0%

Total Count 12 2 3 3 20
% within VARQGOOO3 100.0% 100.0% 100.0% 100.0% 100.0%
YL LT SSSI ETT TE SSIES OSS NRCC ek a 7 ECE eS
184 @ STATISTICS FOR THE SOCIAL SCIENCES

Exercise 6.4
In the table above, a nation’s Civil Liberties score is dependent on the independent
variable Political Rights. Interpret the table. VAROO004 is the Civil Liberties score,
and VARO0003 is the Political Rights score. Only the column percentages are
reported here.

Table E6.5.1 Table of Telephone Lines Per 100 Population by GDP/Capita

Case Processing Summary

Cases

Valid Missing Total

N %o N % N %

VAROO005 * VAROOOO1 20 100.0% 0 O% 20 100.0%

VAR00005 * VARO0005 Cross-tabulation

VAROOOO1

1.06 (2.00 3.00 4.00 Total

VAROO005 1.00 Count 6 1 0 0 Z


% within VAROOOO1 85.7% 33.3% 0% 0% 35.0%

2.00 Count | 2 3 0 6
% within VAROOOO1 14.3% 66.7% 50.0% 0% 30.0%

3.00 Count 0 0 3 4 7
% within VAROOOO1 0% 0% 50.0% 100.0% 35.0%

Total Count Z 3 6 4 20
% within VAROO001 100.0% 100.0% 100.0% 100.0% 100.0%

Exercise 6.5
Interpret the table above. VAROO005 is Telephone Lines Per 100 Population, and
VAROOO1 is GDP/Capita.
Here are the same data for Exercise 6.5 done with SAS. B is Telephone Lines and
A is GDP/Capita. All three percentages were generated here. You would use the
column percentages, the bottom one in each cell.
Constructing and Interpreting Contingency Tables ® 185

Table E6.5.2 Table of Telephone Lines Per 100 People by GDP/Capita

The FREQ Procedure


Table of B by A
B A
Frequency
Percent
Row Pct
Col Pct 1 2 3 4 Total
| 6 1 0 0 a
30.00 5.00 0.00 0.00 35.00
85.71 14.29 0.00 0.00
85.7 | 33.33 0.00 0.00
2 1 2 3 0 6
5.00 10.00 15.00 0.00 30.00
16.67 3338 50.00 0.00
14.29 66.67 50.00 0.00

5 0 0 3 4 7
~ 0.00 0.00 15.00 20.00 35.00
0.00 0.00 42.86 57.14
0.00 0.00 50.00 100.00

Total Z 3 6 4 20
35.00 15.00 30.00 20.00 100.00
eeeee ee ee

Table £6.6 Table of Health Spending as Percentage of GDP by GDP/Capita

The FREQ Procedure


Table of C by A
C A
Frequency
Col Pct 1 2 3 4 Total
ES
1 2 1 1 0 4
28.57 33.35 16.67 0.00 35.00

4 2 | 2 9
2
57.14 66.67 16.67 50.00

] 0 4 Zz Z
3
14.27 0.00 66.67 50.00

‘Total z 3 6 4 20
186 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 6.6
In the above SAS table (only frequencies and column percentages are generated
here), variable C is Health Spending as a Percentage of GDP/Capita. Is there a rela-
tionship? If so, what kind?

Table E6.7 Table of Corruption Perception Index by GDP/Capita

The FREQ Procedure


Table of D by A

D A

Frequency
Col Pct 1 2 3 4 Total

] 4 2 0 0 6
57.14 ~ 66.67 0.00 0.00

2 3 1 0 0 4
42.86 33.33 0.00 0.00

3 0 0 3 0 3
0.00 0.00 50.00 0.00

4 0 0 3 4 7
0.00 0.00 50.00 100.00

Total 7 3 6 4 20

Exercise 6.7
The Corruption Perceptions Index (D) is the row variable. (Remember the higher
the ranking, the lower the perceived corruption.) GDP/Capita is the independent
variable. Interpret the table.

Table E6.8 Table of Corruption Perception Index (dichotomized) by


GDP/Capita (dichotomized)

VAROOOTIO

1.00 2.00 Total

VAROOO1 1 1.00 Count 10 0 10


% within VAROOO10 100.0% 0% 50.0%

2.00 Count 0 10 10
% within VAROOO10 0% 100.0% 50.0%

Total Count 10 10 20
% within VAROOO10 100.0% 100.0% 100.0%
Constructing and Interpreting Contingency Tables » 187

Exercise 6.8
Here, Corruption and Perception and GDP/Capita have been dichotomized by
combining High and Medium (1, the new High) and combining Low and Very Low
(2, the new Low). VARO0011 is the new Corruption Perception Index, and
VAROO010 is the new GDP/Capita. Using SPSS this time, compare these results to
Table E6.7. What has happened?

Table E6.9 Table of Political Corruption Perception Index (dichotomized) by


GDP/Capita (dichotomized), Controlling for the Civil Liberties
Index (VAROO004)

VARO0011 * VARO0010 * VARO0004 Cross-tabulation

VAROOO10

VAROOOO4 1.00 2.00 Total

1.00 VAROO011 1.00 Count 8 8


% within VAROOO1O 100.0% 100.0%

Total Count 8 8
% within VAROOO10 100.0% 100.0%

2.00 VAROOO11 1.00 Count 2 0 2


% within VAROOO1O 100.0% 0% 33.3%

2.00 Count 0 4 4
% within VAROOO10 0% 100.0% 66.7%

Total — Count 2 4 6
% within VAROOO1O 100.0% 100.0% 100.0%

3.00 VAROOO11 2.00 Count 3 &


% within VAROOO1 0 100.0% 100.0%

Total Count 3 3
% within VAROOO10 100.0% 100.0%

4.00 VAROOO11 2.00 Count 3 3


% within VAROOO10 100.0% 100.0%

Total Count 3 3
% within VAROOO10 100.0% 100.0%
cc
lL...

Exercise 6.9
Here Civil Liberties is used as a control variable. The table between Corruption
Perceptions and GDP/Capita is reproduced for each category of Civil Liberties:
(1) High through (4) Very Low. See if you can interpret these results. Hint: For the
_ highest Civil Liberties Category (1), there are eight countries—all both high in
Corruption and in GDP.
STATISTICS FOR THE SOCIAL SCIENCES

Table £6.10 Table of Corruption Perceptions Index (B) [dichotomized]


by GDP/Capita (A) [dichotomized] and Controlling for the
Political Rights Index C

The FREQ Procedure


Table of B by A
B A

Frequency
Col Pct 2 Total

1 10 0 10
100.00 0.00
2 0 10 10
0.00 100.00

Total 10 10 20

Table 1 of B by A
Controlling for C =1

B A

Frequency
Col Pct 1 2 Total

1 10 0 10
100.00 0.00
2 0 2 2
0.00 100.00

Total 10 10 12

Table 2 of B by A
Controlling for C =2

B A

Frequency
Col Pct / 2 Total

1 0 0 0
0.00
2 0 2 2
100.00

Total 0 2 5
Constructing and Interpreting Contingency Tables » 189

The FREQ Procedure


Table 3 of B by A

Controlling for C =3
B A

Frequency
] 9 oe
Col Pct
0 a
| 0

0.00
3 é
ye 0

100.00

3 4
Total 0

Table 4 of Bby A
Controlling for C =4

B A

Frequency
1 5 el
Col Pct

0.00

100.00

Total 0 3 3

Exercise 6.10
Here is Corruption Perceptions by GDP/Capita (both dichotomized), run using SAS
and controlling for Political Rights. The first table is without a control and resembles
Table E6.8 in content. Each table that follows controls for a category of Political
Rights. Interpret these results.
ce emeemmeemermeeamiael
VY KEY CONCEPTS ¥
‘SRR Sa EP Ese aaa Rr ENOL ALAN OGLE LOONEY EERE I EE LET IIIT

tests of statistical null hypothesis/H, one-way analysis


significance alternative hypothesis/ of variance/one-way
inferential research hypothesis/H, ANOVA
statistics/inductive fallacy of affirming the .05 level of significance
Statistics consequent probabilities
descriptive statistics chi-square test Type I error/alpha error
population/sampling sample statistics nondirectional alternative
universe population parameters hypothesis/two-tailed
sample population mean/mu/u alternative hypothesis/
drawing a sample population standard two-tailed test
sampling bias/ deviation/lowercase directional alternative
biased sample sigma/O hypothesis/one-tailed
random sample one-sample tests alternative hypothesis/
simple random one-sample z test one-tailed test
sample one-sample ¢ test degrees of freedom
sampling error two-sample ¢ test Type I error/beta error
LEELA IIE GE SEALE ERODED SELLE ON ELLEN ETL ELLE LIER EGE LENE DIETER NEEL ELIS RLITE LENE INEST EEE LENO
CHAPTER

Statistical Inference
and Tests of Significance

BodJEROEOGUE M

We nowLane dealing anda a new aree The see isee asaga


scientists (indeed, any scientists), we want to generalize about groups
that are too large for us to study. I want to study juvenile crime, but I can’t
interview every juvenile criminal in the world, let alone in my own country,
or even the city where I live. I want to study married couples, but I can’t
interview every married couple in town, or even in my own neighborhood.
I want to study the psychology of heart patients, but... . So many subjects
to study, so little time!
So I study some juvenile offenders, but not all; some married couples,
but not all; some coronary patients, but not all. Yet it is really the a// about
whom I want to generalize. How safe is it for me to study some but conclude
ee my pacijesywould eppiy ti toolen
NNER ARR HANS ANNIE EEEILO SNE OTEYEE EERE EI ITER IES SIESCSO

191
192. STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION

We now turn our attention to the issue of sampling and the procedures used
to determine the likelihood that data obtained from a sample will reflect
the group from which the sample was selected. In other words, to general-
ize about a very large group—a reference group, all autistic children, or
a large legislative body, for example—we must study a subset, or sample, of
the whole group. How certain can we be in concluding that what we found
in studying the sample really applies to the whole group? To deal with this
problem, we use inferential statistics and tests of statistical significance.
In this chapter, we will provide a road map to four chapters that will
follow. We will discuss the nature of sampling and the logic of testing for
statistical significance. The simplest of these tests, the one-sample zZ test,
is used as a model for presenting the steps performed in all such tests.
Theoretical considerations are postponed until the next and subsequent
chapters. We present major problem formats in this chapter, and the tests
that match each format are mentioned and referenced to the specific
chapters in which they are presented in detail.
A word of caution: Many new ideas are presented here, and careful
reading is required. The logical process that underlies a test of significance
is likely to seem strange to a newcomer. With new ideas to grasp and a new
way of thinking to be mastered, it may take you a while to be comfortable
with this new material. But take heart. Once you know this chapter, most of
what comes in the next four chapters should be fairly easy to grasp. The
“number crunching” will become more complex, but the type of reasoning
learned here will remain the same.

WHAT IS STATISTICAL INFERENCE?

To begin our discussion, let us look at Table 7.1. If the hypothesis is that there
is a relationship between attitude on the death penalty for top drug dealers
and attitude on intervention in Latin America to halt drug production, then it
appears that the hypothesis has been verified. Supporters of the death penalty
predominate 3:1 among supporters of intervention, whereas the ratio is
reversed (3:1 death penalty opponents) among those against intervention. Up
to this point in this text, that is the conclusion we would be expected to make.
We now add a new question. About whom are we generalizing? Certainly
for 30 of the 40 people studied here, the hypothesis is verified, but are
we trying to generalize only about that group of 40 people? Assume they
are college students. Perhaps we want to generalize about all students at
that particular college. Or perhaps we want to generalize about all college
students in the country, or even all college students in the world. Often, we
Statistical Inference and Tests of Significance ($33

want to reach a conclusion for more than the people actually studied.
But how safe are we in concluding that what is true for the 40 students
of Table 7.1 is also true for all students at that college? How safe are we in
generalizing beyond that college to the country? The world?
The techniques that help us to answer this question about generalizing to
a larger group are known as tests of statistical significance, and the body
of knowledge that deals with such tests of significance is called inferential
Statistics or sometimes inductive statistics. Up to this point we have only
been dealing with descriptive statistics, where frequency distributions or
relationships between variables are described. Now we turn our attention to
inferring that what is true for a group we have actually studied is true for the
larger group about which we want to generalize. We are inducing—going
from the specific to the general—that what is true for the subjects studied, the
40 students, is true for all students in the college, the country, or possibly the
planet.
_—emmesemennnmenttannnnainrtnnnnsnee
Tests of statistical significance |Techniques that help us to generalize to a
larger group.

Inferential statistics or inductive statistics The body of knowledge that deals with
tests of significance.

Descriptive statistics Statistics in which frequency distributions or relationships


between variables are described.

Table 7.1 Attitude on Drug Intervention

Attitude on Death Penalty For Against Total

Supports 15 5 20
Opposes > 15 20

Total 20 20 40

Let us begin by introducing some terms used in studies such as the pre-
ceding, where a smaller group is selected specifically to reflect a larger group.
The group about which we want to generalize is called the population or,
less often, the sampling universe. (We call this group the population even
though it may not necessarily be a population in a demographic sense, such
as all American citizens or all residents of the Bronx.) The population consists
of all the members about whom we wish to generalize, for example, all felons,
all attorneys, all divorced women, or all Methodists. The key word is al.
194 STATISTICS FOR THE SOCIAL SCIENCES

TL
NT TS

Population or sampling universe The group about which we want to generalize.


ielpi ne teases eae
From the population, we select the smaller group that we will study—
the sample. We call the selection of the subjects to be in that sample
drawing a sample, in the same sense as drawing cards in a card game
such as poker. How we go about drawing the sample determines whether
we can use a test of significance to help us in our generalization to the
population.

Sample The smaller group from the population that is selected to be studied.

Drawing a sample The selection of subjects to be in a sample.

One way a sample may not accurately reflect the population from
which that sample was drawn is known as sampling bias; that is, a biased
sample could have been drawn. In sampling bias, it is the mechanism for
selecting the sample that causes the sample to be not representative of
the population as a whole. For example, for the study shown in Table 7.1,
suppose we had recruited students by giving priority to recovered drug
addicts. This could possibly result in a group with a disproportionate
number of “hard-liners” on both issues. Thus, the resulting sample would be
not representative of the distribution of opinion in the college as a whole.

Sampling bias or biased sample The mechanism for selecting the sample that causes
the sample to be not representative of the population as a whole.

Sometimes, a biased sample is intentionally selected since the


researcher seeks a particular outcome. For example, suppose I want to
prove that “Paingo” is the best-selling headache medicine in town. I inter-
view shoppers leaving the drug store, ask them if they had purchased
headache medicine, and, if so, what brand. Of the 20 people who said they
had bought headache medicine, suppose 15 said they bought Paingo. I
could then prepare a television commercial claiming that 75% (15 out of 20)
of purchasers prefer Paingo. However, suppose I knew in advance that the
only headache remedies sold by that particular pharmacy were Paingo and
a new medicine for which there was not yet any name-brand recognition by
the consumer. My sample would be intentionally biased.
Usually, there is not such an ill intention by the researcher, but something
not perceived causes the sample to be biased. Imagine the same scenario and
the same results—three out of four customers prefer Paingo—but unknown
Statistical Inference and Tests of Significance > IDs

to me, the drug store had Paingo on sale that day, at half price! No wonder
so
many people preferred Paingo.
In American politics, a classic example of sampling bias took place in
1936 when the Literary Digest polled an unusually large sample and then
predicted that Republican Alf Landon would win the election by nearly 60%.
In the election, the landslide went not to Landon, who only won 38% of the
vote, but to his opponent, Franklin D. Roosevelt. It turned out that the
Literary Digest had selected its sample from lists of automobile owners and
from the phone book. This was during the Great Depression, when only the
relatively well heeled could afford automobiles or telephones, and these
groups at the time were staunchly pro-Republican. The E.D.R. supporters
came from the far larger group of those less well off and less likely to own
what at that time were luxuries.
We must bear in mind that if a sample is biased, there is 7o statistical
technique that can make the sample representative of the population.
Accordingly, every technique we will discuss assumes, among other things,
that the sample is representative, not biased. If there is reason to believe
that our sample is biased, no test of significance exists for turning our sow’s
ear into a silk purse.

RANDOM SAMPLES

If a sample is not biased, it is a random sample. In fact, statistical tests


assume a simple random sample, one in which each member of the pop-
ulation has an equal chance or probability of being included in the sample.
Drawing a simple random sample assumes a selection process not unlike an
honest lottery. Suppose we list all the students in our college alphabetically.
We then assign each student a number beginning with 1 and continuing until
all, say, 3000 students have been included. We might then get a rotating drum
containing 3000 marbles of equal size and weight, numbered consecutively
from 1 to 3000. We turn the drum many times to thoroughly mix the marbles
and have a volunteer select a marble. The volunteer takes out the marble,
writes down its number on a list, and returns the marble to the drum. The
drum is turned again, and the process is repeated until our list contains as
many numbers as the size sample we want to select. (Assume for simplicity
that no marble is selected twice and that the drum is sturdy, so that nobody
loses their marbles.) To get a sample of 40 people, we would take the list of
40 numbers back to the alphabetized list of all students in our population,
write down the name adjacent to each number on our list, and use this list as
our simple random sample. Every student in the college would have had an
equal likelihood of being included in the sample.
196 STATISTICS FOR THE SOCIAL SCIENCES

el

Random sample A sample that is not biased.

Simple random sample A sample drawn in such a way that every member of the
population has an equal likelihood of being included in the sample.
snannnaensemerennnenneenmntaiasainammmmnannanmttitiininnnnneentsnmmnnt

This tedious technique is unnecessary today, as we may instead use a


pregenerated table of random numbers or a random number-generating
computer program. All honest sampling techniques used by legitimate
researchers and polling organizations strive to produce unbiased samples.
These techniques rarely generate perfect random samples, but they come
close enough that we may use them and assume randomization.
Suppose we have used all means to ensure that our sample of
40 students was randomly drawn. We no longer have to worry about
sampling bias, but another nemesis awaits: sampling error. The term
error as used here does not mean that a mistake was made in selecting the
sample. Rather, error here means deviation from what actually exists in the
population. This deviation can be accounted for by the laws of chance or
probability. For example, imagine a population that is 50% female. Using a
proper randomization method, we draw a sample that just by chance is 60%
female. We mistakenly conclude that the population is 60% female. In a
few samples, this can happen by chance. In a somewhat smaller number of
samples, we might get 70% women. In an again smaller number of samples,
we might get 80% women. There is even some likelihood (“once in a blue
moon”) of getting all women in our sample and mistakenly concluding that
the college population is entirely female. Similarly, any other characteristic
of a sample may deviate from the population due to sampling error—
religion, income, party preference, fraternity membership, or, in the case of
our original problem, attitudes on the two variables.
OMANI EDELSTEIN ARAN ANEE AEEN RELL LEED SELES NIA NIE OLESEN OOOELE RETEST EEE EE NO EEE EIEN ERE ONT NEON EEO

Sampling error A deviation from what actually exists in the population not
associated with sampling bias but still existing, even though the sample was
randomly drawn.

Suppose in the college of our study, there is 70 relationship between


attitudes on the death penalty and intervention, and yet as a result of
sampling error, we drew a random sample in which there was such a rela-
tionship. A test of significance would tell us how likely it is that this could
have occurred by chance, given random selection of the sample. It would
tell us the odds of concluding, based on our sample, that in the population,
our two variables are related when in fact they are not. Note that such tests
Statistical Inference and Tests of Significance 197

do not tell us for any given sample whether that sample accurately reflects
the population. The tests only tell us the probability that this may be the
case. Thus, we will be living with uncertainty from here on.
We begin by stating a null hypothesis, symbolized by H,, with the
H meaning hypothesis and the subscript meaning zero, for null. The nature of
a null hypothesis varies from problem to problem. In the case of our problem
in Table 7.1, it is a statement that in the population of all students at this col-
lege, there is no relationship between attitudes toward the death penalty and
drug intervention. It states that the two variables are independent of one
another and that the relationship appearing in the table is solely the result of
sampling error (chance) and does not reflect a relationship in the population.

Null hypothesis A statement postulating that in the population, the means of two
or more groups are the same or, in the case of two variables in a cross-tabulation,
that in the population, the two variables are unrelated.
8

H,: In the population, there is no relationship between the variables.

If H, is true, we would be wrong to conclude that the variables are


related in the population. What if H, is probably false? Then we might con-
clude one of several possible alternative hypotheses (sometimes called
research hypotheses), the simplest of which is that in the population, the
two variables are related. We call the alternative hypothesis H, (we usually
formulate only one alternative hypothesis, but we could have others—H,,
H,, etc.). Please note that if one examines actual published scholarly
research, the research hypotheses are often explicitly stated, but the null
hypotheses are often not mentioned at all. Nevertheless, the logic of research
assumes their existence.
Ne aBeeeeessai8N

Alternative hypothesis or research hypothesis Statement that indicates that in


the population, the means of two or more groups differ or, in the case of a
cross-tabulation, the two variables are related.
TPOLNEESS

If H, is true, then the relationship in the table is not due to sampling


error but is actually reflective of a relationship in the population. Note that,
as a researcher, you are really trying to prove H,, but a principle of logic
known as the fallacy of affirming the consequent suggests that we can
“prove” H, only indirectly by showing that /) is probably untrue. If H, is false,
then H, is the only logical conclusion. If H), which states that there is no rela-
tionship in the population, is false, then the only other general conclusion is
H,: In the population, there is a relationship. We set up H, hoping to disprove
198 @ STATISTICS FOR THE SOCIAL SCIENCES

it; it is a “straw man” whom we hope to knock down. We will return to this
problem in Chapter 12 and make use of a test of significance known as the
chi-square test to see if we can reject the null hypothesis. That test will be
appropriate whenever two variables are presented in a cross-tabulation such
as Table 7.1, regardless of the level of measurement of those variables.
MOMURSSAMOM
MME SAMO EL EE SEA MNS,

Fallacy of affirming the consequent A principle of logic that suggests that the only
way to “prove” the alternative or research hypothesis is to demonstrate that the null
hypothesis is untrue.

Chi-square test A test of significance for two variables in a cross-tabulation.

COMPARING MEANS

Let us now consider a situation where we are comparing two groups’ means,
such as the means of two classes that have taken a common examination. If
we treat both groups as populations, our conclusions are made simply by
examining and comparing the two means. Thus, either the mean for Class 1 =
the. mean for Class:2, or the mean. for Glass.1.4 the mean for Class. 2: Ui the
latter case is correct, either the mean for Class 1 > the mean for Class 2, or the
mean for Class 1 < the mean for Class 2. We simply compare the two numbers
to reach our conclusion. However, when one or both of the means comes
from random samples rather than populations, the possibility of sampling
error emerges, and we need to perform a test of significance on the data.
Before discussing this further, we need to introduce some new terms.
To differentiate data from a sample from data from the population, we call
the statistics computed from sample data sample statistics and those from
the population data population parameters. We designate sample statis-
tics with the same notation we have been using all along; that is,

xX is a sample’s mean.

sis asample’s standard deviation.

s’ is a sample’s variance.
n is a sample’s size.
EON ORE TEMS EEE ON COC EH TRUER CT IE ENNIS SENN PES TEN TEE TE TEE NENT NESSIE HSI IY NEE OTN TE HTN ESO

Sample statistics Information computed from sample data.

Population parameters — Information computed from population data.


==Ee ESSERE NSIS SANE URS SER ESS
Statistical Inference and Tests of Significance > 199

For sample statistics, we also continue to use the formulas we have learned
thus far.

= 2:
n

x.—a?
X
v= PaileeOr (definitional formula)

or

XS deat —_ ———__

v= 2 £ (computational formula)
n
For population parameters, we use lowercase Greek letters, often with
subscripts to differentiate them.

Greek letter mu. is a population’s mean.


Greek letter sigma. 0° is a population’s standard deviation.
O° is a population’s variance.

N is a population’s size.

Greek letter mu (y) Represents a population’s mean.

Greek letter sigma (o) Represents a population’s standard deviation.

Our formulas then become

cee ia

(TO 2_ Lien)?
x (definitional formula)

Or

XG. ee — BILD
o* = 2 ea (computational formula)

Now, in a problem calling for a test of significance, we may no longer


know both of the population means. In fact, we may know neither of them.
200 STATISTICS FOR THE SOCIAL SCIENCES

Yet, on the basis of sample means (theX s), we want to make a generalization
about the respective population means (the ps).
Our null hypothesis is that there is no difference between the two
population means.

FL): by = be

If the test of significance enables us to reject H, we will conclude that

A: Ub, # ML,

Let us examine an example where this kind of null and alternative


hypothesis would be formulated. Suppose we wanted to compare Ohio’s
Democratic members of the House of Representatives to their Republican
colleagues, in terms of liberalism. Here liberalism is a score assigned to
each representative’s voting record by Americans for Democratic Action
(ADA), a liberal organization. For Ohio’s 21 representatives in the 101st
Congress, we have ADA scores for 20 people (all 11 Democrats and 9
of the 10 Republicans).' We will treat the 11 Democrats and the 9
Republicans as the populations of all Ohio Democratic and all Ohio
Republican representatives, respectively. We calculate the two population
means and compare them:

democrats =
oni
83
ee)

Republicans =17.. bo2

Obviously, the two us are unequal to each other; in fact, the « for the
Democrats is much higher (more liberal) than the w for the Republicans.
Note that we now have evidence to verify the hypothesis that there is a rela-
tionship between party identity and ideology, such that Democrats tend to
be more liberal than Republicans. Since we have compared two population
means, no sampling is involved, there is no sampling error, and no test of
significance is needed.’

Comparing a Sample Mean to a


Population Mean or Other Value

Suppose, however, we knew the yu for the Democrats but did vot know
it for the Republicans. Suppose instead that we had access to a random
sample of size m= 3 of Ohio’s Republican representatives and that the
sample’s mean was 23.33.
Statistical Inference and Tests of Significance » 201

X Reps = 2IO9

lene = 02:18 Mreys = unknown

We formulate a null hypothesis:


~

i, 0° democrats “ HRepublicans

Now, if the null hypothesis is ¢rze, Lpsststons = pendent = OO: Lon ENG TAC
that our sample x for the Republicans, 23.33, is different from 83.18 would
be attributed to sampling error. In other words, due to random chance, we
drew a sample from a population whose mean was 83.18 and got a sample
mean of 23.33,
If, on the other hand, we could reject H,, we could conclude instead that
our sample probably did mot come from a population whose mean was
83.18, but rather from a population whose mean differed from 83.18.

HI, MDdemocrats - HRepublicans

We would then refine our conclusion even more by noticing that since 23.33
is less than 83.18, Mpenuplicans 8 probably less than 83.18, and Republicans are
probably less liberal than the Democrats.
The tests of significance that we would perform to see if H, could be
rejected are called one-sample tests since we are comparing data from
one group’s sample to another null-hypothesized value—in this case, data
from another group’s population. Specifically, depending on the informa-
tion given to us, we would do either a one-sample z test or a one-sample
t test. We will examine the former test later in this chapter and discuss both
tests in the next chapter.

One-sample tests Tests that compare data from a sample to similar data in a
population.

One-sample z test A test of significance that can be performed when we know the
population’s standard deviation as well as its mean.

One-sample t test A test that can be performed when we know the population’s
mean but not its standard deviation.
202 STATISTICS FOR THE SOCIAL SCIENCES

Comparing a Sample Mean to Another Sample Mean

Suppose both geoupticans 2G Hpemocrats Were unknown. We would then


have to draw random samples from both groups and compare the two
sample means. Assume the same figures for the Republican sample as used
above. In addition, we draw a sample from the Democrat population, say,
a sample of 2 = 4 Democrats with a calculated sample mean of 82.50.

X= 82.50
“~Dems
Mie ee oo
Leoems = UNKNOWN reps = unknown

Our H, and H, will remain as before:

las 0° Democrats a MRepublicans

FL: Democrats - MRepublicans

Since our comparison is now between two sample means, we could


perform a test known as the two-sample ¢ test (discussed in Chapter 9). We
would follow a line of reasoning similar to that used for the one-sample tests
discussed earlier. An alternative to the two-sample ¢ test is the one-way
analysis of variance or (one-way ANOVA, for short). Generally, though,
ANOVA is more commonly used when there are more than two groups to be
compared.

Two-sample f test A t test that compares two sample means, rather than one
sample’s mean to another population’s mean.

One-way analysis of variance (one-way ANOVA) A test in which two or more


sample means may be compared simultaneously.

Comparing More Than Two Sample Means

In our Ohio example, there are only two parties to compare. What if
there were more than two parties? For instance, suppose a similar study had
been contemplated for the Canadian House of Commons. Our null hypoth-
esis might look like this:

Jad 0° MBio« Québécois ~ Kv iberals = conservatives = Upp

As mentioned, one-way analysis of variance would be used to compare


these means (see Chapter 10),
Statistical Inference and Tests of Significance j» 203

The data situations, tests of significance, and chapters of this text


pertaining to them are summarized below:

Summary
Data Situation Test of Significance Chapter
One sample mean versus> One-sample z test reas)
a population mean Or one-sample ¢ test 8
One sample mean versus Two-sample ¢ test 2
another sample mean
Comparing several One-way analysis 10
sample means of variance
Comparing two variables Chi-square test for ZZ
in a cross-tabulation contingency

THE TEST STATISTIC


In the case of the z and ¢ tests, the size of the z or ¢ generated is, in part,
a function of the distance between the means being compared. In this
one-sample case, this is the difference between the sample mean and the
population mean. (In other one-sample cases, the sample mean might be
compared to some designated value other than a population mean.) In the
two-sample case, it is the distance between the two sample means.
Let us return to the Ohio delegation problem, where

Mdems = 83. 18

reps= unknown

Recall that we drew a sample of 7 = 3 Republicans and got a sample mean


Of X po, = 23.33. If we also knew the standard deviation of the population,
Opems WE Would be ina position to do a one-sample z test. (We will see in
the next chapter that if we do not know o,,.,,,, we have to estimate it from
our sample data, and we would do a one-sample ¢ test instead of a z test.)
Please keep in mind that what follows is an overview of testing for
significance as we will be doing in the following four chapters. Thus, in
chapters to come, the origin and meaning of each formula will be explained.
The formula for the one-sample Zz test is

X=
o//n
204 << STATISTICS FOR THE SOCIAL SCIENCES

For our specific problem, the formula becomes

a XReps — //Dems
Opems//N

Once we know 6,,,,,, we can calculate z (even though we do not yet know
what to do with that information). We find that 6,,,,.Dems = 10.5, and accordingly,

__ XReps — Dems
- Opems//n

Phe
refe Renk oS ts)
10.5//3
= 59.05
10.5/1.732
_ —59.85(1.732)
in 10.5
103.06
10.5

= —9.872
The negative sign on z is due to the fact that X,..,, is less than py... If X > M,
then z would be positive. For purposes of deciding statistical significance,
we will use the absolute value of z,thus disregarding its sign!
For now, let us leave our calculated z of —9.872 and look at what happens
to z as the distance between X and wu increases (and thus the numerator of the
z formula increases). We will first imagine a case where xX and uw are the same.

lye 65.18 — 95,18 0 0(1.732) 0


if xk=yl; 2 = eee ee ees
10.5//3 LOS 1.752 10.5 105
Now let us look at what happens to z as X gets farther away from wu. We
will reduce x by increments of 10 units at a time and see how it affects the
absolute value of z.

If ¥ = 83.18 |x -p| =0 lz|=0


73.18 10 1.649
63.18 20 3.299
53.18 30 4,948
43.18 40 6.598
33.18 50 8.247
23.18 60 9.897
eres) 70 11.546
2.18 80 13.196
Statistical Inference and Tests of Significance jp» 205

If the null hypothesis is true and y,,,.,. = Lgeps, We Would expect the mean
of a sample of Republicans to be very close to the population mean for
Republicans. Ideally, they would be the same. If 11... = Preps ANd fpens =X pens)
then logically X,.., — pems = 9, and z will be zero. Ideally, if the null hypoth-
CSISMSHthuUeeza—1 0)
However, even if the null hypothesis is true, it still is gzite likely that
the Republican sample mean will be slightly different from the Republican
population mean due to sampling error. Thus, X,., — Mpems Could often be
slightly different from zero, and z could be slightly different from zero as
well. Note that as the gap between X,.,, and p.,,, grows, the likelihood that
the null hypothesis is true shrinks. It is always possible that the null hypoth-
esis is true, no matter how far X,.., is from [u,,,,, and thus how large a z we
get. But as the gap between xX,.,,. and fp,,,, increases and thus z increases,
the likelihood that the null hypothesis is true decreases.
In our Ohio example, nearly 60 points separate our X,,,, Of 23.33
and our Up.,,, Of 83.18. Our z of —9.872 is quite large in absolute value, as
compared to z scores generally encountered. It is true that given a true
null hypothesis, we could get an X that is different from p and a z that
large due to sampling error. But it is so improbable that in this case, an
explanation other than sampling error would be much more plausible in
accounting for our large z. The alternative explanation is that the popu-
lation from which the Republican sample was drawn has a mean different
from the mean for the Democrat population. In short, it is our a/terna-
tive hypothesis:

A; MDemocrats a MRepublicans

At what point do we decide that the z obtained is large enough to reject


the null hypothesis? While many factors may enter into this decision in
applied research, in the confines of the classroom, we use a common
convention called the .05 level of significance. If 5 (or fewer) out of 100
random samples (i.e., 1 out of 20 or fewer) drawn from a population where
the null hypothesis is true yield a z value equal to or greater than the one
we obtained, we will reject the null hypothesis and conclude instead our
alternative hypothesis, H,. In doing so, we run a risk not to exceed a .05 pro-
portion (5%) that really H, is true and we are mistakenly rejecting it. the .05
proportion is really the probability that we are falsely rejecting a true null
hypothesis.

.05 level of significance The probability level classically used by statisticians for
determining that a null hypothesis may be rejected.
SALORORaLLI ORAS BE OSES IYEOLA MIELE EEE EEE MILL NEN ANNE ESSE AERA,
enntaon eSB SISA HOES ON NETIEEIEEE
206 << STATISTICS FOR THE SOCIAL SCIENCES

PROBABILITIES

Let us briefly examine the nature of probabilities, which can be defined


as proportions that reflect the likelihood of a particular outcome occurring.
An easy-to-understand analogy is the batting average used in baseball.
Imagine a simplified game in which walks, bean balls, and other anomalies
are eliminated so that a batter stepping up to the plate faces one of two
possible outcomes. The batter either gets a base hit (be it a single, double,
triple, or home run) or does not get a base hit (mighty Casey strikes Out).
For any given “at bat,” we do not know whether the player will get a hit
or not, but we can state the probability of a hit occurring by examining the
batter’s past record. Suppose the batter has come to bat 100 times and has
had 30 base hits. The probability that he or she will get a hit when coming
to bat is the number of accumulated base hits divided by the total accumu-
lated at bats. In this case,

base hits 30
Batting Average (probabilityof a base hit) at bars 100 3

Probabilities Proportions that reflect the likelihood of a particular outcome


occurring.

Thus, there is a .300 probability (a 30% chance) of the batter getting a


base hit. Our player is batting 300 (announcers often don’t use the decimal).
Although this tells us our player’s likelihood of getting a hit is a little less
than one out of every three at bats, we know for any given at bat, a batter
will still either get a hit or not—one cannot make .300 of a base hit. None-
theless, we would prefer to send a batter to the plate whose batting average
is .300 rather than to bring up someone with a .150 batting average!
By using the .05 probability level as a decision point in our test of
significance, in effect we are requiring a “batting average” of .950 before we
reject the null hypothesis. At that point, the risk of error in rejecting the null
hypothesis is .05 or 5%. We will continue our discussion of probability in the
next chapter.

DECISION MAKING

Recall that in our Ohio problem, we had obtained a z of —9.872. To know


whether we can reject the null hypothesis, we compare the absolute value
of the obtained z, 9.872, to a series of critical values of z. These critical
Statistical Inference and Tests of Significance j» 207

values (whose origins we will discuss in the next chapter) are values of
z for differing levels of significance (probabilities of error in rejecting 1)
beginning with the crucial .05 level.

Probability (Level of Significance) Critical Value of z


Sa oe 1.96
O1 Zo
001 3.29

Here, the probability is that of falsely rejecting a true null hypothesis. This is
also referred to as a Type I error or an alpha error.’

Type | error or alpha error The probability of falsely rejecting a true null
hypothesis.

The procedure is to compare the absolute value of the calculated, or


obtained, z to the critical values above. Since the .05 level is our decision
point, the following two possible overall outcomes exist.

A. |zobtained [19 6e
In this case, ourz is less than Z,irica at the .05 level. We cannot reject H,.
We say that the difference (between X,,,, aNd Mpems) iS Not statistically
significant and imply that the difference between X and p is the result of
sampling error.

B. ee real ? 1.96.

In this case, z equals or exceeds Z,,,;.4) at the .05 level. We reject Hy


(which means we accept H,). We say that the difference between 0 and
us is statistically significant.

In our Ohio problem, z = -9.872, |z| = 9.872, and 9.872 > 1.96. Thus,
we reject the H, that there is no difference in liberalism between Democrats
and Republicans in Ohio’s congressional delegation. The difference
between 23.33 and 83.18 is probably not due to sampling error; instead, it
probably reflects a real difference between population means.
Now, if |Zpainea! < 1-96, we have completed our task. If, on the other
hand, |Zjraineal 2 1.96, we need to take a further step. Remembering that we
could be making a mistake in rejecting H,, we need to report to the reader
the likelihood, or odds, that we are making an error. Recall that if Z pained
had exactly equaled 1.96, the probability of error would be exactly .05. Thus,
5 out of 100 similar-sized random samples drawn from a population where
208 < STATISTICS FOR THE SOCIAL SCIENCES

H, is true would generate zs of 1.96 or more. If our sample had been 1 of


those 5 samples, we would make a mistake in rejecting H).
Suppose we had obtained a z of exactly 2.58—1 out of 100 samples
will yield a z that large due to sampling error. If we then reject H,, we are
falsely rejecting a true null hypothesis. The probability of this happening
is .01. We might even have obtained a z of 3.29—the result of sampling
error in 1 out of 1000 samples, with a probability of .001.

Note: The larger the z obtained, the smaller the probability of making
such an error.

When using the z test, we use these three levels as benchmarks for
reporting the probability of a Type I error. (Other tests may use additional
probability levels below .065—more on that later.) If we had done our Zz test
with a packaged computer program, it would have told us the exact proba-
bility of alpha. Having done this by hand, we instead report the probability
by the range into which it falls, as follows:

A. If |Zirainea| < 1-96 and we cannot reject H,, we report 7o probability.


By Plz ee = Le put obtained
|< 2.58, we rejectH, and reportp < .05;
that is, the probability is less than .05 and (by convention) greater
than .01. Sop <.05 tells us that the probability is between .05 and .01.

CH 12 eweal 2 Do DU lease oo, wre TeeClr. aug rept gy< 07.


Here,p is less than .01 but greater than .001.

D. If |Z,praineal 2 3-29, we reject H, and reportp < .001. We generally stop


the process of reporting probabilities ofz at the .001 level.

In our Ohio example:

lz
|“ obtained =O 0 1a L9G, reject,
= 9.872 > 2.58
SO. Bia > Si295 so p < .001

Review

Remember that for the one-sample z test, we are given «and o for one
population. For the random sample drawn from the other population, we
know the sample’s size 7 and its mean X. Once we have this information, we
use the formulaz= (X — 1) / (6 / Yn to find z, sometimes referred to as SA cisia!
We compare |Zpraineal tO Zeritic At the .05 level (i.e., 1.96). If |z| is less
than 1.96, we cannot reject H,. If |z| 2 1.96, we reject H,. If so, we compare
|z| to the critical values of z to determine the probability of a Type I error.
Statistical Inference and Tests of Significance » 209

Examples

Suppose we have a scale measuring environmental activism, ranging


from 0 to 100 (most active), designed to tap individual attitudes and behavy-
ior concerning the environment. We wish to test the hypothesis that, in the
population, the level of education and environmental awareness are posi-
tively related. Suppose we know that for the population of all those who
have graduated from high school but not college, the mean score is 50 with
a standard deviation of 10. We then draw a random sample of 100 college
graduates, administer the same survey, and calculate the sample means that
are presented below. The null and alternative hypotheses are as follows:

lel 0° righ school — Meollege

AL: Mrigh school a Meollege

Suppose our sample mean was 51.7. We have all the information needed
for a one-sample z test.

eNotes, pei 50 5171050 ae lO


le
~ Ohs//2 - 10/4/00 10/10 1
Since 1.70 < 1.96, we cannot reject H,. We cannot conclude that the mean
environmental activism scores of the two populations differ.
What if our sample mean was 52?


ee! Z 2
ey
i ee—

10/100 10/10 1

Since 2.00 > 1.96, we can reject H,, but since 2.00 < 2.58, we can only report
PaO:
What if X...
coll
= 53?

a
ele ee ene
=
eer
10/100. 10/10 1

< .01.
Since 3.00 > 1.96, we can reject H,. Then, 3.00 > 2.58 but 3.00 < 3.29, sop
What if X.., = 54?

— 4
= a ates: i = — = 4.00
10//100 10/10 1
> 3.29,
Since 4.00 >:1.96, we can reject H,. Then, 4.00 > 2.58 and 4.00
sop < .001.
210 STATISTICS FOR THE SOCIAL SCIENCES

DIRECTIONAL VERSUS NONDIRECTIONAL ALTERNATIVE


HYPOTHESES (ONE-TAILED VERSUS TWO-TAILED TESTS)
In the problems we have done so far, our alternative hypotheses have looked
like this:
Ai: by, F by

Specifically,

AL: Udems of Lreps and A: Mi. a Keo

Note that when we use the inequality symbol #, we allow for two possible
conditions.

bby > My or by < My


We do not specify which of the two conditions is likely but instead collect
our data and calculate z.
In the Ohio example, we were able to reject H, with p< .001, thus
concluding [poms # Meeps: At this point, after all the facts were in, it would
be reasonable to say that since X,,,. = 23.33, which is /ess than the Lyon, OF
83.18, empirically Republican representatives have lower ADA scores than
Democrats in Ohio. Ultimately, then, we are really concluding Up.n, > Mreps»
which is even more specific than our initial H/,.
When our H, has the format pw, # “,, we refer to it as a nondirectional
alternative hypothesis—it does not specify which direction, “, > mw, or
[L, < (4), Will ultimately be correct. For reasons to be explained in the next
chapter, this is often referred to as a two-tailed alternative hypothesis
or a two-tailed test of significance. We will see that the number of tails
noted in the expression may not always be correct, making this terminology
ambiguous. Nevertheless, these expressions are so much a part of statistics
tradition that they are constantly used.

Nondirectional alternative hypothesis (two-tailed alternative hypothesis or


two-tailed test of significance) An alternative hypothesis that does not specify the
directionality (i.e., which mean will ultimately be the larger).

The nondirectional (two-tailed) H, is the more traditional and more


conservative format for the alternative hypothesis, but it is often replaced
by a directional or one-tailed 7, (or a one-tailed test). To formulate a
directional alternative hypothesis, one must be able to discard, prior to col-
lecting the data, one of the two possible directions implicit in the H,, either
Statistical Inference and Tests of Significance j» 211

ML, > fk, OF LW, < ,. When one of the directions is discarded,
the remaining
direction becomes the alternative hypothesis.
pene
Directional alternative hypothesis (one-tailed alternative hypothesis; one-tailed
test
of significance) An alternative hypothesis that does specify which mean will be the
larger one.
LL
Oe SAeS Heme t

In our Ohio example, the nondirectional H, of tp... # reps Subsumes


two possibilities: either

Upems - Hreps

Or

Upems < LReps

What if, in advance of examining the data, we reviewed the logic of these
two possible directions. This assumes that we have prior knowledge about
Democrats and Republicans. Given such knowledge, is it more logical to
assume Democrats are more liberal than Republicans or less liberal than
Republicans? With the exception of Southern Democrats, all evidence sug-
gests that Democrats are more liberal than Republicans. In fact, recent
Republican campaign strategies have been aimed at reinforcing just such an
impression. That being the case, can we eliminate in advance the possibility
Of Myems < Hreps? If so, we could formulate a directional H, as follows:

fe Fe dems me Reps

(Obviously, if we had evidence that Republicans are the more liberal of the
two, our AH, would be Mpeg < Mreps:)
As we will see in the next chapter, picking a directional H, gives the
advantage of making it easier to reject the null hypothesis. The directional
critical value is always less than the nondirectional one. We can see this in
the following sets of critical values.

Critical Value of z

Probability One-Tailed Two-Tailed


(Level of Significance) (Directional H, ) (Nondirectional H, )

.05 1.65 1.96


O1 2.33 2.58
001 3.09 3:29
Suppose that in the high schooi versus college example, we had prior
evidence showing that college graduates were more environmentally active
212 < STATISTICS FOR THE SOCIAL SCIENCES

than high school-only graduates, so that p,, > H.,, Was an illogical assumption
to make. Accordingly, we form a directional H, as follows:

Ai: Mrs < Moot

Note that the computation of z is exactly the same as when H, was


nondirectional. In the case where X...,coll = 51.70,

Be Se
sO = é
er 1.70
os oe dee. lOy10 ail
Before, since 1.70 < 1.96, we could not reject H,. Now, however, by
making a directionality assumption in H,, we may make use of the lower
one-tailed critical values.

Za OS 165 Reject H,
170 <n2535 pews

We are now able to reject H,, whereas without the directionality assumption,
we could not. For that reason, directional alternative hypotheses are widely
used.
Despite their wide usage, there are major risks associated with one-
tailed alternative hypotheses. For example, is there really prior evidence on
which to make a directionality assumption? A researcher may give little, if
any, rationale for the direction chosen in the assumption. Moreover, with
the use of modern multivariate techniques, there may be dozens of variables
being manipulated at once. The more data, the less likely that each pair of
means or pair of variables has been systematically examined to justify the
directionality of each possible alternative hypothesis. Thus, be wary of
conclusions from one-tailed tests!

Setting the Level of Significance

The .05 level usually determines the significance decision because of


custom and tradition. It represents a compromise between the risk of falsely
rejecting a true null hypothesis and the risk of falsely retaining an untrue
null hypothesis. As we increase the former risk, we decrease the latter risk,
and conversely. At times, researchers may decide to make it harder to reject
H,, by basing the decision on a higher level of significance, such as .01. For
example, to be even more certain that a drug being tested was superior to
another and that the test results were not due to the effects of sampling, a
researcher might use the .01 level. In most social science applications, the
need for a higher level of significance is rarely encountered. In fact, there
Statistical Inference and Tests of Significance ® 213

is pressure to go the other way and use, in formal academic application, a


lower level than .05 for rejecting the null hypothesis! There is ongoing
discussion and debate on this matter.
Even now, there are circumstances where we may make it easier to
reject the null hypothesis. For instance, Z,,,,,.., at the .10 level (directional) is
1.282. Suppose we obtain a z= 1.5. By our earlier standards, we need at least
1.65 to reject H,. But, we can only say that the difference is significant at
the .10 level if it was stated earlier (with appropriate justification) that the
basis for significance decisions in the study was the .10 level. Otherwise, we
could be accused of trying to sneak in as significant something that would
normally be considered not statistically significant.
As a matter of practicality, particularly in nonacademic settings, there are
circumstances where a lower level of significance makes sense. To understand
this, recall that earlier we noted that the size of z is based in part on the dif-
ference between X and yw. However, the size of z is also determined by the size
of the sample. The larger the 7, the larger the z. For example,

7 — OO

22 50. 1.70V100 1.70 x 10


= = 1.70
10/./100 10 10

but if 72 = 1000

Sot 50 Ly 0V 1000" UVOGIC22) 99-797 4a, 5.375


S104 1000 tn, 10 10 10

Using the two-tailed critical values, the z where m = 100 is not significant,
whereas the z where 7 = 1000 is significant, p < .001. Given a big enough ”,
even trivial differences become statistically significant.
By contrast, suppose our sample 7 was lower than 100, say, 25.

_ 51.70-50 2 1.70/25 _ 1.706) — 0.85


KOH oe 10 10

Here, z drops from 1.70 to 0.85 even though X and w were the same.
Keep in mind that survey research costs money, and a major factor in
the expense is the size of the sample to be interviewed. For example, sup-
pose some local government wants you to do a study of some public service
,
delivery but is only willing to pay you $2000. Because of the dollar limitation
214 <4 STATISTICS FOR THE SOCIAL SCIENCES

you discover that your sample size cannot exceed 50. Based on earlier
studies, you feel that you need at least 100 people to get results significant
at the .05 level. Since you cannot get funding for 7 = 100, you tell the con-
tracting officer that you will do the survey if the city will accept a lower level
of significance, say, .10. The contracting officer—even in the unlikely event
that he or she knows what you are talking about—may be willing to accept
your suggestion, just to stay within the budget. Finally, you may wish to do
a pilot study on a small sample as part of what will eventually be applied to
a larger sample. Here you are interested in eliminating ambiguities from
your survey document. You accept the .10 level of significance, knowing that
in the final study, 7 will be large enough for the .05 level to be used.
In more and more published research, you are likely to encounter a trend
where the probabilities are simply stated, with no statement as to whether
the results are statistically significant. Then it is up to you, the reader, to
examine each probability and make your own conclusion about significance.
We have now learned that two factors play a role in determining the
magnitude of the z obtained: the size of the difference between X and y and
the size of the sample, 7. We will examine this subject again in Chapter 9.

Degrees of Freedom

In the case of the one-sample z test, the critical values of z remain


constant. For example, at the .05 level, nondirectional H,, the critical value
of z is always 1.96. With other tests of significance, however, the critical
values will vary from one problem to another depending on something
called degrees of freedom, d.f, or just df for short. In these cases, to
determine what critical values to use, we must first find the degrees of free-
dom. Beyond z, each variation of t, F or chi-square will also carry a formula
for finding df We will not attempt to define degrees of freedom at this time
other than to say that in these tests, dfis related to the sizes of the samples
being used or, in the case of chi-square, the sizes of the tables.

Degrees of freedom An additional piece of information needed for tests where


critical values vary with the problem and may be functions of such things as
sample size.

CONCLUSION

Steps in Significance Testing

In this chapter, we have introduced the logic of testing for statistical


significance as well as the procedures followed for any test of significance.
Statistical Inference and Tests of Significance jp 215

Using the one-sample z test as a reference, let us review these steps in their
proper logical order.

1. Before any data are examined, formulate a null hypothesis and an


alternative hypothesis. If you are going to use a directional or one-tailed
,, make sure that there is convincing evidence to justify making the
directionality assumption needed for your alternative hypothesis. Also,
justify any decision to use a level of significance other than .05 for bas-
ing your conclusion about whether the obtained statistic is statistically
significant.

2. Once the data are collected, make sure that you have or can obtain
the information needed to perform the test of significance you have chosen.
For the one-sample z test, you will need the mean and standard deviation
for one population and the size and mean of the sample you are comparing
to that population. In short, you need pu, 6, 2, and Xx.

3. Calculate the test statistic.

4. Locate the appropriate critical values, which you will compare to


the statistic you obtained. Make sure that the values are appropriate for the
H, you formulated—directional or nondirectional. In cases other than the
z test, you will first have to determine the appropriate degrees of freedom
in order to find the appropriate critical values.

5. Compare the statistic that you obtained (calculated) to the appro-


priate critical value at the .05 level (unless you have chosen another level for
your decision).

6. If your obtained value is less than the appropriate critical value, you
cannot reject H,. You cannot say that the difference is statistically significant.

7. If your obtained value equals or exceeds the appropriate critical


value, you can reject H, and say that the difference is statistically significant.

8. If a significant difference is found, you still may be making a Type I,


or alpha, error—that is, falsely rejecting a true null hypothesis. Report the
probability that this is the case by comparing your obtained statistic to crit-
ical values at progressively higher levels of significance. The level that p is
less than is the level of the last critical value that your obtained statistic
exceeded.

9. Go back to your null and alternative hypotheses and, based on the


results of your significance test, restate your conclusion.
216 @ STATISTICS FOR THE SOCIAL SCIENCES

Chapter 7: Summary of Major Formulas

Population Parameters

The Mean

ae
aN
L
The Variance The Variance
Definitional Computational

RS Oe
> d~e&-p)?
anes 2 Bee ees
N N

- Sample Statistics
The Mean

eee ae
|
The Variance The Variance
Definitional Computational

Sox? a Ca:
2 it
a =
V1

The z Test of Statistical Significance


xXx— ph
o//n
Statistical Inference and Tests of Significance j» 217

EXERCISES
Exercise 7.1
Write a null hypothesis and a nondirectional alternative hypothesis for each of the
following.
Example: Scottish voters are more supportive of the Labour Party than English
voters. Assume that the variable is a measure of pro-Labour attitudes. Thus,

HH, Uscots = english

Ay: Uscots # Henglish

Women differ from men on their attitudes toward abortion.


Urban residents differ from rural residents on the issue of gun control.
The French support state subsidies to their farmers more than do Americans.
Israeli Arabs differ from Israeli Jews on the issue of trading land for peace.
B&B
mM
WH Liberals in the United States support affirmative action programs more than do
conservatives.
6. Australians and New Zealanders differ in their support for a U.S. nuclear pres-
ence intheir region.
7. French Canadians are less culturally assimilated to the larger English-speaking
culture than are the Cajuns of Louisiana.
8. Soviet military officers differed from Soviet civilians in their level of support
for perestroika.
9. Californians tend to support pro-environmental legislation more than do New
Yorkers.
10. Police officers support the death penalty more than do convicted murderers.

Exercise 7.2
Recall the tests of significance discussed in this chapter:

One-sample z test or t test: Compares the mean of a sample to the mean of a


population.
Two-sample t test: Compares the mean of one sample to the mean of another
sample.
One-way analysis of variance: Compares the means of several samples,
generally more than two.
Chi-square: Tests the significance within a cross-tablulation.

Examine each of the following problems and indicate which of the above tests
is most appropriate for that problem.
218 @ STATISTICS FOR THE SOCIAL SCIENCES

1. H,: Members of the House District of Columbia Committee (assuming they are
selected at random) are significantly younger than the overall membership of
the House of Representatives.
For the House District Committee, the mean age is 35.
For the House of Representatives, the mean age is 45.
For the House of Representatives, the standard deviation is 14.
The House District Committee has 11 members.

2. Hy: In the population from which a random sample was drawn, respondents’
dogmatism scores are unrelated to socioeconomic status (SES) classification.
Assume: Dogmatism is an interval scale.

SES
High Medium Low
Dogmatism 1 a 10
Scores Zz > 2
‘ 2 8 8
N equals 18 3 8 10
4 4 10
1 6 )

3. H,: In the population (all countries), there is an association between the


number of political parties and that nation’s level of political development.

Level of Development
Traditional Modern
(Underdeveloped) (Developed)
Mean number of political
parties per country 3.0 2.8
Sample size 15.0 25.0
Sample variance 4-228 6.0

4. A random sample of 25 convicted felons is studied to see if there is a relation-


ship between a convict’s income level and the type of sentence received. Treat
both variables as nominal dichotomies.

Sentence
Income Level Fine Jail Term | Total
High 5 10 | is
Low 0 10 | 10
Total dye 26 | 28
Statistical Inference and Tests of Significance » 219

Hj: In the population, participation level in professional organizations is


unrelated to occupation type. Assume: Organizational Participation is an
interval scale (0 = minimum to 10 = maximum)

Occupation Category
, Professionals Nonprofessionals
Respondents’ 10 3
organizational S) 5
participation 9 4
scores: 6 2
8 0
& |
1
1
ms ~
Sample means: 8.5 Ze.
Sample sizes: 6 9

H,: In the population from which a random sample was drawn, type of
occupation is unrelated to job satisfaction.

Type of Occupation
Job Satisfaction Professional White-Collar Blue-Collar Farmer Total
High 35 20 5 5 65
Low # 10 toe NSP 3s
Total 540 3 30 30 2 © 190
. Hg: |n the population, participation in professional organizations is unrelated to
type of occupation. Assume: Organizational Participation is an interval scale
(0 = minimum to 10 = maximum activity).

Occupation Category
Lawyers Doctors Other Professionals
Respondents’ 10 4 7
organizational 2] 5 5
participation 8 > 8
scores: 10 a 9
10 {| 4

9 6
Sample means: 93 4.4 6.5
Sample sizes: 6 . 6
220 STATISTICS FOR THE SOCIAL SCIENCES

8. Assume that for the entire population of the Irish Republic, the mean age is 30
years. A random sample of 15 members of the Dail (the Lower House of
Parliament) yields a mean age of 45 and a standard deviation of 15.
H,: There is no difference in mean age between the population of the Irish
Republic and the members of the Dail.

Exercise 7.3
Each of the following examples looks like a problem calling for a one-sample z test
or t test. In each case, however, a flaw in the logic of the research design makes a
test of significance moot. For each example, identify that flaw.

- At Breezewood Junior High School, there are a total of 50 eighth-graders in two


government classes, taught by Mr. Jones and Mrs. Smith, respectively. At the
end of the quarter, both classes take a common civics exam.
H,: In this instance, in fact, Mrs. Smith’s class scores higher than did Mr. Jones’s
class.

Class
Mr. Jones Mrs. Smith
Mean exam score 86 87
Class size 25 25
Class standard deviation 12 14

. Suppose for the U.S. population as a whole, it had been determined that
the mean assertiveness score was 50 on a scale ranging from 0 (/east) to 100
(most). A researcher wishing to generalize about the Dayton metropolitan
area’s residents’ assertiveness characteristics selects as the experimental group
a random sample of 25 jet pilots stationed at a nearby Air Force base. The
researcher then administers the assertiveness test to them and obtains a sample
mean of 70 and a sample standard deviation of 12.
Hy: » for the Dayton metropolitan area equals ut for the United States.

. Atacertain university, the mean undergraduate grade point average is 2.9, with
a standard deviation of 1.2, for the entire undergraduate student body. For the
50 political science majors in the honors program, a random sample of n = 20
has a mean of 3.4,
H,: The mean of the population of political science majors is higher than the
mean for all undergraduates.

. At the same university as in Example 3, the mean grade point average for all
psychology majors is 3.3. Is the difference statistically significant when these
majors are compared to all undergraduates at the university?
Statistical Inference and Tests of Significance j» 221

Exercise 7.4
Calculate the one-sample z test of significance for each of the following. Just
calculate z.
G10 = 12 O=5 p= 25
2. 6 = 150 uu = 100 6=25 n= 100
a, XS] w= 2.8 621.2 n= 36
4. X= 32 LL = 30 o=10 n=49
D>. £=2.6 u=3.0 o=1.4 n= 64

Exercise 7.5
Using the nondirectional (two-tailed) critical values of z, examine each obtained z
below. Reach a conclusion about statistical significance and, if significant at least
at the .05 level, state the probability of alpha, using a “p <” statement.

Ngee = 3.0)
24.7) 2-00
3.62 == 2.65
4. z=-1.90
5. Z =—2.58
OF 1 75
ho f=-3.10
8. z=—1.50
O 2= 42.33
10.2 2 50

Exercise 7.6
Repeat Exercise 7.5 using the directional (one-tailed) critical values of z.

Exercise 7.7
Five of the examples in Exercise 7.1 could be directional alternative hypotheses.
Identify them and write the appropriate directional alternative hypotheses.

Exercise 7.8
Formulate the null hypotheses and the most appropriate alternative hypotheses
(either directional or nondirectional, as you think appropriate) for each of the
following. If H, is directional, justify the direction.

1. Managers will differ from their employees in their support for business interests.
2. Southern senators will differ from senators in general in terms of their support
for conservative policies.
3. A group of people who have watched a video with a decidedly pacifistic message
will have a different attitude toward the use of military force than people in general.
222 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 7.9
Following are the data pertaining to the three items in Exercise 7.8. _
For each, follow Steps 2-9 as outlined in the Steps in Significance Testing on
page 215. (You have already done Step 1 in Exercise 7.8.)

1. For all managers, on a scale ranging from 0 to 100, the mean business support
score is 36.1 with a standard deviation of 11.1. For a random sample of 9
employees, the mean business support score is 81.8.
2. For all U.S. senators, the mean conservatism score is 56.2 with a standard
deviation of 32.4. For a random sample of seven Southern senators, the mean
score is 61.9.
3. Fora pacifism scale ranging from 0 (low pacifism) to 10 (high pacifism), the mean
pacifism score is hypothesized to be 3.0 with a standard deviation of 0.5. For a
random sample of 25 people who watch the video, the sample mean is 3.2.

NOTES
1. From H. Stanley, and R. Niemi, Vital Statistics on American Politics, 3rd ed.
(Washington, D.C.: CQ Press, 1992).
Following are the data from which the means were calculated:

Democrats’ ADA Scores Republicans’ ADA Scores


70 35
dS 20
75 3
90 >
90 0
95 15
95 15
70 30
95 30
90
70
> = 915 > Y= 155

O15 ie)
pens = —— = $3.18 LReps =
1]

There was no score for one Republican.


Statistical Inference and Tests of Significance » 2s

2. Some people argue that such tests may be applied even in comparing two
populations. In such a case, there is no sampling, but the populations could differ
only due to chance randomizations in nature. In this text, however, we will ignore
this debate and exclude tests comparing two or more populations.
3. We report the alpha or Type I error’s probability whenever we reject the null
hypothesis. Note that if we do vot reject the null hypothesis, we could be making
another type of error: failure to reject a false null hypothesis. This is known as a
Type II or beta error. Its probability, however, is mot reported, even though we
should be aware that it exists. The relationship between Type I and Type II errors is
important and will be revisited in Chapter 9 when we discuss statistical power.
LESSEEOOS

Type Il or beta error Failure to reject a false null hypothesis.


NH SEE LAOS

4. The subscripts obtained or obt., used with Zojpiineg and other tests in
subsequent chapters, is often omitted in statistics tests. We use it here, on a selec-
tive basis, to assist you in learning this material and differentiating the obtained
values from the critical values of the test.
W KEY CONCEPTS ¥

normal! distribution normality assumption confidence intervals for


standard score law of large numbers means and proportions
one-sample z test i test conditional probability
sampling distribution of degrees of freedom addition and
sample means “sigma-hat” (6) multiplication rules
central limit theorem z test for of probability
standard error/standard proportions permutations
error of the mean interval estimation combinations
LTE LOMAS EMS AAI ELLIE LED DIESEL ELAS SLES DL SEEN LER ALLIENSAE LEBEN EINE LEESON LES LEER NEES EOS ODE EDEL EIU OLE NEIL
GIT OTE
CHAPTER

Probability Distributions
and One-Sample z
and ¢ Tests

WY PROLOGUE ¥

This chapter is essentially an extension of the previous chapter, except that


in addition to presenting several new topics, we go back to explain the
theory that underlies the z formula. What are we actually doing when we
calculate 2?
It isn’t absolutely necessary to understand the underlying theory simply
to work these formulas, any more than it is to understand the physics and
chemistry of the cooking process in order to cook a meal. Nevertheless, it
can be useful to understand what you process in order to cook a meal. Now
that you are hungry, let’s get back to statistics. You (or your computer) can
always calculate z, 4 & and so on. Still, it is useful to be able to visualize and
understand what is actually happening when you do these tests.
Cll MES ee s ERUNIWG

pe 225
226 << STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
In the previous chapter, we discussed tests of significance and used the
one-sample z formula to illustrate the entire procedure.

- Ol 7
In this chapter, we turn our attention to the origin of that formula and
explain what is taking place when we use it. It is possible to use any statisti-
cal formula without such an understanding and simply plug in the numbers
as we did in Chapter 7. But if you can visualize what is going on, your under-
standing will be enhanced since essentially the same process takes place no
matter what test of significance is being performed.
The z test of significance is based on a frequency distribution known as a
normal distribution and is applied to a specific normal curve called the sampling
distribution of sample means. When we perform this test, we are actually taking
the given sample statistics and population parameters and locating them on the
sampling distribution of sample means. In fact, all tests of significance do the
same thing, even though their sampling distributions differ from one another.
At the end of this chapter, we will discuss the one-sample ¢ test, and in sub-
sequent chapters, the other commonly used tests of significance will be pre-
sented. We will begin by discussing an even simplerz formula than the one in
the previous chapter and introducing the concept of a normal distribution.

NORMAL DISTRIBUTIONS

Normal distributions are a family of frequency distributions that, when


graphed, often resemble bells. Generally, they are represented in the form
found in Figure 8.1. Such a curve has three major characteristics: (a) it is
unimodal, (b) it is symmetric, and (c) it is asymptotic to the x-axis. This
last characteristic, which becomes very important later on, means that the
tails of the curve get closer and closer to the x-axis but never reach it.
Consequently, no matter how far you get from the mean on the x-axis, there
will always be a tail continuing beyond that point. The tail never ends, at
least not in the mathematical model.
Sinead edassea ahs aineseae andenaeaiadaaeadaceaatae
Normal distributions A family of frequency distributions that, when graphed,
often resemble bells.
iain hese lata aleae ee a ee
The reason for using the term normal distribution is that certain char-
acteristics such as human height, weight, or intelligence graph in frequency
distributions approximating this bell-shaped pattern. The term normal is a
Probability Distributions and One-Sample z and t Tests » 227

Figure 8.1 = The Mathematical Model of the Normal Distribution

Figure 8.2. [Q as a Normal Distribution

bit misleading, however. Not everything in nature is normally distributed, so


nonnormal distributions are not really abnormal.
Let us use measures of intelligence—intelligence quotients or IQs—to
illustrate the normal distribution (see Figure 8.2). IQs are designed to range
from 0 to 200 with a mean of 100. Standard deviations vary by the age of the
subjects but are usually around 13 or 14; for computational ease we will use 10.
Note that, unlike the mathematical model in Figure 8.1, in Figure 8.2 the tails
do end—at IQs of 0 and 200, respectively. This indicates that there are natural
upper and lower limits in the actual measurement tool being utilized. For
example, say it is a test with 200 questions and someone’s IQ is the number of
correct answers. Two geniuses each score 200, yet, if one of the two is really
twice as smart as the other and the IQ test had 400 questions instead of 200,
the one genius would score 400 whereas the other would only get 200.
Suppose Sandra takes this IQ test, and her score, which we designate as
have
x, is 115. We would like to know what proportion of people would likely
higher IQs than Sandra’s and what proportion would have lower IQs. (Sorry,
turns
this procedure cannot tell us how many will have exactly Sandra’s IQ.) It
228 < STATISTICS FOR THE SOCIAL SCIENCES

out that the area under the curve corresponds to the proportion of people
with a particular characteristic. The total area under the curve (1.00 propor-
tion) accounts for all (100%) people. Since a normal curve is symmetric, .50
proportion (50%) of the area under the curve falls below the mean, and .50
proportion falls above the mean. Thus, half of all people should have IQs
below 100, and half should have IQs above it. This proportion of the area also
pertains to the probability of randomly selecting a person with a particular
characteristic. Since .50 proportion of the area of the curve is below the mean,
there is also a .50 probability of randomly selecting a person whose IQ is
below 100. Likewise, there is a .50 probability of randomly selecting someone
whose IQ is greater than 100.
In Figure 8.3, we have added Sandra’s IQ, x= 115. The proportion of
people with an IQ greater than 115 is the shaded area under the curve in the
right tail, fromx = 115 tox = 200. The proportion of people with IQs below 115
is represented by the remaining unshaded area under the curve, from the left
of x = 115 tox =0. Note that the unshaded area has two components, the .50
proportion of IQs less than 100 plus the area under the curve from 100 to 115.
To find these areas, we use a table of areas under the normal curve that
applies to all normal distributions. To use the table, we begin by calculating
what is called a standard score, which is universally designated by the letter
z. (Its relationship to our z test of significance will be explained later.) To
calculate z, we recast the distance from the mean (100) to the value of x we
are studying (Sandra’s IQ of 115), expressed as standard deviation units. The

Figure 8.3

Note: To aid visualization, this figure has mot been drawn to scale.
Probability Distributions and One-Sample z and t Tests » 229

distance from the mean of Sandra’s IQ is x — 4 = 115 — 100 = 15; her IQ is 15


points greater than the mean. To convert that distance into standard deviation
units, we divide by the size of the standard deviation, o = 10, which was given
to us. Since 15/10 = 1.5, we know that Sandra’s IQ is 1.5 standard deviation
units from the mean. Expressing the whole process in a single equation, we get

Standard score A score universally designated by the letter z, in which that


score is expressed in standard deviation units from the mean.

eeOn Le
a =—= its)
oO 10 10

We now go to Table 8.1 and see that each page has three blocks of
figures, and, in turn, each block has three columns: A, B, C. Column A lists
a value of z. Column B shows the area under the curve from the mean out
to that specified value of z (note the graphs above each column). Column C
shows the area in the tail beyond z. Note that the area in column B plus the
area in column C always add to .5000. Also note that as z gets larger, the area
in column B gets larger, and the area in column C gets smaller.
The bigger the z, the smaller the tail.
We find our z of 1.5 at the bottom of the center block on the second
page of Table 8.1. Note that at z = 1.5, the number in column B is .4332, and
the number in column C is .0668. The area in the tail, corresponding to the
proportion of people with IQs greater than 115, is the number in column C,
.0668—only 6.68% have an IQ higher than Sandra’s. To find the proportion
with IQs less than 115, we take the area in column B and add to it .50, the
proportion with IQs below the mean.

IQs between 100 and 115: .4332


IQs between 0 and 100: .5000
Total = .9332

Thus, .9332 proportion of people has IQs below 115. Sandra’s pretty bright!
If Sandra is bright, George is not; his IQ is only 80. Let us find the pro-
portions of area above and below 80 (see Figure 8.4). First we find z.

x—p 80-100 —20 = — 2.00


Seige = Rui ake LO
Here, z is negative since x is less than py. In Table 8.1, we will look for the
to
absolute value of our z (2.00) but use the graphs at the bottom of the page
see that now the shaded areas are to the left of the mean. In the center of the
230 << STATISTICS FOR THE SOCIAL SCIENCES

Table 8.1 Proportions of Area Under Standard Normal Curve

B G A B G. A B G
A
/ an ay

z , ai \ oa JH Ne ez Jh AS ol ah

0.00 .0000 5000 0.30 1179 3821 0.60 2257 2743


0.01 .0040 4960 0.31 mlerall 3783 0.61 2291 .2709
0.02 0080 4920 0.32 IRA Ss: 3745 0.62 2324 .2676

0.03 .0120 4880 0.33 1293 SOM, 0.63 arom! 2643


0.04 .0160 4840 0.34 mlsoul 4669 0.64 .2389 2611

0.05 .0199 4801 035 1368 3632 0.65 2422 WSs


0.06 .0239 A761 0.36 1406 3594 0.66 2454 2546

0.07 .0279 yi 0.37 1443 DoDy, 0.67 .2486 .2514


0.08 0319 4681 0.38 .1480 D20 0.68 ZO 2483
0.09 .0359 4641 0.39 Sy 3483 0.69 2549 2451

0.10 .0398 4602 0.40 1554 5446 0.70 .2580 2420


‘O).1lal 0438 4562 0.41 NES! 3409 0.71 26 2389
0.12 .0478 4522 0.42 1628 So72 0.72 2642 2358
OMS O57 4483 0.43 1664 3336 0.73 .2673 Zoe
0.14 .0557 4443 0.44 .1700 3300 0.74 2704 2296

Onis .0596 4404 0.45 1736 3264 0.75 2/34 .2266


0.16 0636 4364 0.46 lee: 3228 0.76 .2764 2236
ON 0675 4325 0.47 .1808 3192 OF 2794 2206
0.18 .0714 4286 0.48 .1844 156 0.78 2823 Ag
0.19 OVS3: A247 0.49 .1879 Hal 0.79 2852 2148

0.20 .0793 .4207 0.50 OS 3085 0.80 2881 AALS)


0.21 0832 4168 0.51 .1950 3050 0.81 .2910 .2090
0.22 0871 4129 ORS 1985 3015 0.82 2939 2061
0.23 .0910 4090 0.53 2019 2981 0.83 2967 Oss
0.24 0948 4052 0.54 .2054 .2946 0.84 2995 2005

ORS .0987 4013 O)S35; .2088 2912 0.85 ROS) oy


0.26 1026 3974 0.56 aAlys) 2377 0.806 soOsyll 1949
Oy, 1064 3936 Oy Sy 2843 0.87 3078 OZ
0.28 AANGE) 3897 0.58 .2190 .2810 0.88 3106 .1894
0.29 meal 3859 0.59 2224 ATA 0.89 pls) .1867

A B G A B G: A B eG

2 MN L\. 2 ML. 2 MN
Probability Distributions and One-Sample z and t Tests » 231

A B C A B Cc A za CG

bs phew eA eRe
0.90 25 159 1841 1.21 3869 Sil iS Oh .0643
6.91 3186 1814 2g 3888 LIL? 153 4370 .0630
0.92 212 1788 123 3907 .1093 1.54 4382 .0618
0.93 3238 LP TS 1.24 3925 HOS LES)S, 4394 .0606
0.94 3264 1736 1.25 D944 1056 1.56 4406 .0594

095 3289 AeA 26 3962 1038 boys 4418 0582


0.96 3915 1685 L2H 3980 1020 1.58 4429 0571
O97 3340 .1660 28 O09) 1003 15? 4441 0559
0.98 3305 .1635 L28 4015 0985 1.60 4452 .0548
0.99 3389 1611 130. 4032 .0968 1.61 4463 0537

1.00 3413 1587 Eea 4049 0951 1.62 4474 .0526


1.01 3438 1562 62 4066 .0934 1.63 4484 .0516
1.02 3461 153? 133. 4082 0918 1.64 4495 0505
1.03 3485 £515 1.34 4099 0901 1.65 4505 0495
1.04 3508 1492 15 4115 0885 1.66 A515 0485
1.67 4525 0475
1.05 coe! 1469 1.36 4131 .0869
1.06 3554 1446 1.37 4147 0853 1.68 4535 .0465
1.07 ‘opie 1423 1.38 4162 .0838 1.69 4545 0455
1.08 Bee 1401 U59. 4177 0823 1.70 4554 .0446
1.09 O21 ADT 1.40 4192 .0808 17 4504 .0436
1.41 4207 .0793 72 4573 0427
1.10 3643 sleet
js 3065 Bile oe) 142 4222 0778 il73) 4582 .0418
1.12 3686 1314 1.43 4236 .0764 1.74 BSED 0409
115 3708 2g 7 1.44 4251 0749 gs 4599 0401
1.14 2729 27k 1.45 4205 0735 1.76 .4608 .0392
1.77 .4616 .0384
115 3749 1251 1.46 4279 0721
1.16 jo L230 1.47 4292 .0708 1.78 4625 0378
Ag 2 3790 1210 1.48 4306 0694 LD .4033 .0367
1.18 .3810 1190 1.49 ‘4319 .0681 1.80 4041 10359
1.50 4332 .0668 1.81 4649 0351
lame) 3830 1170
1.20 3849 iS io A345 .0655 1.82 4656 .0344

Me a 2 MS
(Continued)
232 << STATISTICS FOR THE SOCIAL SCIENCES

Table 8.1 (Continued)

A G A B G ‘a B C

vA IVa 2S CF Wee ae Zi A ; KYL is ~ 4

1.83 4664 .0336 Pails, .4834 .0166 2.43 4925 0075


1.84 4671 .0329 2.14 4838 .0162 2.44 4927 .0073
1.85 .4678 0322 Zale) 4842 .0158 2.45 4929 0071
1.86 .4686 .0314 2.16 4846 0154 2.46 4931 0069
1.87 4693 0307 BMF 4850 0150 2.47 4932 .0068

1.88 4699 .0301 2S 4854 .0146 2.48 4934 0066


1239 .4706 .0294 MNS) 4857 0143 2.49 4936 .0064
1.90 4713 0287 220) 4861 .0139 2.50 4938 0062
1.91 4719 0281 Pagal 4864 .0136 Zeal 4940 .0060
ILS 4726 .0274 Dgaph 4868 .0132 sya 4941 0059

1°93, 4732 .0268 Cn, 4871 0129 ZD9 4943 .0057


1.94 4738 0262 2.24 4875 0125 2.54 4945 0055
Gs 4744 0256 PAs .4878 0122 2.55 4946 0054
1.96 4750 0250 2.26 4881 (US, 2.56 4948 .0052
Ly 4756 0244 ZOee, 4884 0116 2.57 4949 0051

1.98 4761 10239 7X8) .4887 0113 2.58 4951 0049


iILDy, 4767 0269) BAS 4890 .0110 Bese 4952 .0048
2.00 4772 .0228 2.30 4893 .0107 2.60 4953 0047
2.01 4778 0222 Peeiik .4896 0104 2.61 LBS) 0045
2.02 4783 0217; 2Eoe, 4898 .0102 2.62 4956 0044

2.03 .4788 O22 Ze, .4901 .0099 2.03 4957 0043


2.04 4793 0207 2.34 4904 0096 2.64 4959 0041
2.05 4798 0202 2.29 4906 0094 2.65 4960 0040
2.06 4803 LOND 7 2.36 4909 0091 2.66 4901 0039
PAO .4808 .0192 ZO 4911 0089 2.6); .4962 0038

2.08 .4812 .0188 2.38 4913 0087 2.08 4903 .0037


2.09 .4817 .0183 239 4916 0084 2.69 4964 .0036
ZO 4821 0179 2.40 4918 0082 OR IAO) 4965 0035
Pa MAL 4826 0174 2.41 .4920 0080 Pal 4906 0034
Pe\\72 4830 0170 2.42 4922 .0078 ZZ 4967 .0033

A
wee B A
ey: A
ae:
PAK a + SV ese aS
Probability Distributions and One-Sample z and tTests ® 233

CARN ANT AA
2.73 4968 0032 2.94 4984 0016 S15 A992 0008
2.74 4969 0031 2.95 4984 .0016 516 Hoo2 .0008
213 4970 .0030 2.96 4985 0015 Sak 4992 0008
206 4971 0029 2.97 4985 0015 518 4993 .0007
ey, 4972 0028 2.98 4986 0014 byw) 4995 0007
2.78 4973 0027 2.99 4986 0014 3.20 4993 0007
2:79 4974 .0026 3.00 4987 0013 Dia 4993 0007
2.80 4974 .0026 5.08 4987 0013 B22 4994 .0006
2.81 4975 0025 3.02 4987 0013 p48) 4994 .0006
2.82 4976 .0024 105 4988 0012 eae 4994 0006
2.83 4977 .0023 3.04 4988 0012 B25) 4994 .0006
2.84 4977 0023 e105 4989 0011 3.30 4995 0005
2.85 4978 0022 3.06 4989 OO11 oo 4996 0004
2.86 4979 0021 S07 4989 0011 3.40 A997 0003
ap eS 4979 0021 3.08 4990 0010 3.45 4997 0003
2.88 4980 0020 O09) 4990 0010 3.50 4998 0002

io 4981 0019 3.10 4990 0010 3.60 4998 0002


2.90 4981 0019 3.11 A991 0009 5070) Bese) 0001
PES 4982 0018 OZ 4991 0009 3.80 4999 0001
2.92 4982 .0018 Bild 4991 0009 3:20 4999 .0000
2:93: 4983 0017 3.14 4992 0008 4.00 4999 0000

A B G A A

: VGN OOhrs a ae
SOURCE: Abridged from R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricultural and
Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint of Pearson Education.

left-hand block of the third page of the table, you will find z = 2.00. The col-
umn B area is .4772, and the column C area is .0228. Accordingly, only .0228
proportion of people has IQs below George’s. To find the other proportion,
add the column B figure to the .50 whose IQs exceed the mean: .4772 + .5000
= .9772 proportion. Poor George! Nearly 98% of all IQs exceed his.
Now we can solve a mystery: the source of the critical values of z used
in the previous chapter. Although a much more detailed table of areas under
the normal curve is needed to find all the critical values of z presented in
Chapter 7, we can find approximate values using Table 8.1.
234 @ STATISTICS FOR THE SOCIAL SCIENCES

Figure 8.4

|
|
|
|
|
|
|
|

ale
0 x = 80 p= 100 200

z=-2.0

Note: To aid visualization, this figure has mot been drawn to scale.

We simply specify a particular tail area, find it in column C, and read the
corresponding z score from column A.
For example, to find the z value for the one-tailed .05 level, we see that
in column C, the two closest approximations are .0505 (z = 1.64) and .0495
(z = 1.65). Actually, the mean of the two z values, 1.645, is the true critical
value, but when we round to two decimal places, we get 1.65. Likewise, the
closest tail area to .01 is actually .0099, and itsz is 2.33. For the .001 level, we
find .0010 occurring three times, where z is 3.08, 3.09, and 3.10. If this table
had used more decimal places, we would see that 3.09 would be our best-
fitting value of z.
For a two-tailed test, we need two tails whose areas, added together,
equal the probability level desired. We take one half of the probability level
as the area to locate in column C. For instance, at p= .05, we need two equal
tails whose areas add to .05, so .05/2 = .025. Finding .0250 in column C, we
see that z = 1.96. (Unfortunately, Table 8.1 is not complete enough for us to
find the other two-tailed values of z.)
Figures 8.5 and 8.6 summarize the relationships between z values and
probabilities for one and two tails, respectively.

THE ONE-SAMPLE z TEST FOR STATISTICAL SIGNIFICANCE

The formula we used in the previous chapter to test for statistical signifi-
cance, z= (x— [1)/(G/V), is really a reworking of the formula we have
Probability Distributions and One-Sample z and t Tests jp» 235

Figure 8.5 — Critical Values of z—One Tail

f=

|
|
|

Area =
Probability = .05
lt —_> X=

~I

Area =
Probability = .01
—> x=

—— |
Hi JdBh3)
p= \
A
|
I
i}
|
|
i}
i}
|

Area =
Probability = .001
+

z=3.09

been using in this chapter, except that it is applied to a specific type of


frequency distribution known as the sampling distribution of sample
means. This sampling distribution is the frequency distribution that
would be obtained from calculating the means of all theoretically possible
samples of a designated size that could be drawn from a given popula-
tion. To illustrate this definition, let us imagine a population of 5 people
with scores ranging from 1 to 5. Using a sample size of 7 = 3, what are
the different combinations of scores that we would obtain in selecting all
236 STATISTICS FOR THE SOCIAL SCIENCES

Figure 8.6 — Critical Values of z—Two Tails

f=

Total Tail Area


= Probability
.025 + .025
ol.05

Area = .025 IpArea = .025

a P
EE >
z=-1.96 Z=1.96

i} i}
i) |
i] 1 Total Tail Area
= Probability
= .005 + .005
if 1
=.01
1 i}
| i}

: InArea = .005

_ ue
Sa |
z=-—2.58 Z=2 50)

f= ;
A
Total Tail Area
= Probability
; = .0005 + .0005
Area= | \ == 4010)
0005 |
1 1 [ Area = .0005

K +
x=

z=-3.29 Fh foie)

possible samples of 3 people out of the original 5? If we consider the order


of selection of the people (e.g., 5-4-3 is one sample, 4-5-3 is another, and
3-4-5 yet another), there are actually 60 possible samples that could be
drawn from this population. If we disregard the order of selection, there
are Only 10 possible combinations of scores.
Probability Distributions and One-Sample z and t Tests j 237

Sampling distribution of sample means The frequency distribution that would


be obtained from calculating the means of all theoretically possible samples of
a designated size that could be drawn from a given population.

For the | Selected Samples of 3 People That Could Be


Population; x = | ~ Drawn From the Population; x =
Dr ein lito co do: teeny Ys
Beh ean OA A AY) eA 14
Bo logs ed ole Le) Ox) £5 3
eal 2 2 Dive
hy a 1 ‘en oh te oa
We will work with these 10 samples of 3 people with the 3 scores shown
in the columns. For each of these samples, we can calculate the mean.

Os eX = XN OS 26 Ne i Oo = I =

Se a ae: 5 4 4 4 3
1 ee an aa ee CE oe ae a
Ole tel ene e 1 2 1 jG
Ne Var ae ees (0 ee 8 9 8 Tene
Boe 20 3 85 8a 3.0 27 16 ay ee ee

The frequency distribution of all these possible sample means is as follows:

x= f=

4.0 1
ee 1 We graph this frequency distribution,
ele) 2 which will approximate the sampling
3.0 2 distribution of sample means, as a
Zo Z histogram, as shown in Figure 8.7.
75) 1
2.0 1

resem-
Note that the histogram in Figure 8.7 has a pattern that begins to
ic. In fact, if
ble a normal curve in the sense that it is unimodal and symmetr
itself normally
either the population from which the samples are drawn is
along the
distributed along the variable x and/or self-normally distributed
populat ion are sufficie ntly
variable x and/or the samples drawn from that
a normal dis-
large, the sampling distribution of sample means will also be
to us.
tribution. This characteristic will prove to be very useful
238 < STATISTICS FOR THE SOCIAL SCIENCES

Figure 8.7. The Sampling Distribution of Sample Means Obtained From


10 Samples of Size 7 = 3 From a Specified Population
NN —$———— ee

The formal statement of these characteristics comes from what is


known as the central limit theorem and a related theorem known as the law
of large numbers.

THE CENTRAL LIMIT THEOREM

According to the central limit theorem, if repeated random samples of


size m are drawn from a population that is normally distributed along
some variable x, having a mean wu and a standard deviation o, then the
sampling distribution of all theoretically possible sample means will be a
normal distribution having a mean uw and a standard deviation o//7.'

Central limit theorem If repeated random samples of size n are drawn from a
population that is normally distributed along some variable x, having a mean p
and a standard deviation o, then the sampling distribution of all theoretically
possible sample means will be a normal distribution having a mean p and a
standard deviation o/./n.

If, for a particular population, some variable (&) is normally distributed


and we draw a series of samples of a predetermined size (m) from that
population, the central limit theorem tells us that
Probability Distributions and One-Sample z and t Tests » 239

1, The sampling distribution of sample means will be a normal distribution.

2. The mean of the sampling distribution of sample means, the mean of


all the sample means (designated xX), will be equal to y, the mean of
the population from which the samples were originally drawn.

3. The standard deviation of the sampling distribution of sample means


will be equal to o/\/7, the standard deviation of the population from
which the samples were drawn divided by the square root of the size
of the samples that we were drawing. This standard deviation of our
sampling distribution, o/\7, is also called the standard error of the
mean or, more often, just the standard error, and is sometimes
designated with the symbol o,.

Standard error of the mean or the standard error The standard deviation of the
sampling distribution, designated with the symbol O,.

To illustrate: Suppose your state or province mandates a series


of competency examinations in reading, math, and so on, to be taken by all
schoolchildren in selected grades. Assume that on the math competency
exam given to all ninth-graders, the mean, pz, is 70, and the standard devia-
tion, 6, is 20. We wish to compare these results to those we would have
found if we had studied random samples of ninth-graders rather than the
whole population. How would the sample means be distributed? Assume we
select random samples of size 7 = 100.
According to the central limit theorem, the sampling distribution of
sample means would be a normal curve with a mean X = uw = 70 and a stan-
dard deviation (the standard error of the mean) 0/7 = 20/100 = 20/10 =
2.0. This is graphed in Figure 8.8.
The implications of Figure 8.8 are immense since they suggest that the
overwhelming majority of theoretically possible sample means are going to
fall very close to the original population’s mean. There are tails to the curve
in Figure 8.8 above and below the mean, and they are asymptotic to the
x-axis, but they are so tiny that they are barely perceivable.
In fact, we know the following about normal distributions: 68.27% of the
area under the normal curve (thus 68.27% of all sample means) falls between
—o-and UL + 6,, that is, one standard deviation above and below the mean.
Since in this particular example, the standard deviation of the sampling distri-
bution (the standard error of the mean) is 2.0, we are observing that 68.27%
of all sample means will lie between 68 and 72. We also know that 95.45% of
the area under the curve (95.45% of the sample means) falls between pt — 20;
and w+ 20,,x? and 99.73% of the area falls between pt — 30; and Mt + 30;.
240 @ STATISTICS FOR THE SOCIAL SCIENCES

In this particular example:

Figure 8.8 — The Actual Appearance of a Sampling Distribution of Sample Means for
Samples 7 = 100 Drawn From a Population Where = 70 and 6 = 20

*II I
OP Ose 0 0 80 90 100

X=

Area Under the


Normal Curve = kange From Range in
Percentage of Mean in Terms of
All Possible the Terms of Competency
Sample Means Standard Errors Exam Scores
68.27% Aol G0 68-72
95.45 Wa 2a. 66-74
Wiehe wt 3 0; 64-76
With 99.73% of sample means falling between 64 and 76, the remaining
0.27% of all sample means must fall below 64 or above 76. Out of 1000
Probability Distributions and One-Sample z and t Tests » 241

samples, only 2 or 3 would have means below 64 or above 76—an extremely


improbable, but statistically possible, event.
The central limit theorem becomes most useful to us when we are given,
as before, w and o for a population and data about one random sample pre-
sumably drawn from that population. In that case, we are asking how likely it
is that, from the given population, we could draw a random sample whose X
differs from « by as much as we observe. If the likelihood is low, we might
better conclude that the x reflects a population with a mean other than the uw
of the population from which we initially assumed that the sample was drawn.
This brings us to the kind of problem presented here and in the previ-
ous chapter. Suppose in our competency exam example, we have a random
sample of 100 ninth-graders who had been enrolled in a 6-week-long course
to prepare for this examination. This sample’s mean score is 73. Thus, H):
Man = course: ASSuming advance data on which to make a directionality
assumption, we could write

A: May) < Hcourse

We calculate z using the formula from Chapter 7 and compare z, obtained to the
critical values of z.

Be a napa ee 5 ee eesti
ofa 20/V100 20/10 2
Since 1.50 < 1.65, we cannot reject H,. The course appears to have been
unsuccessful.
With this formula, we are finding our sample’s X on the x-axis of the
sampling distribution of sample means, finding the distance from that x to
the mean of the sampling distribution, and converting that distance into
standard deviation units (standard scores) based on the standard deviation
of the sampling distribution. To see how this works, let us start by convert-
ing our simple z formula from symbols to words.

a
Hoy
Oo

(1) | The value of the variable | (2) | The mean of our |


| whose distance from the | - | frequency distribution |
| mean we wish to find | |

(3) | The standard deviation of | |


| our frequency distribution |
242 << STATISTICS FOR THE SOCIAL SCIENCES

Remember that the frequency distribution is the sampling distribution of


sample means. Now note the following:

1. The value of the variable whose distance from the mean (of the
sampling distribution) we wish to find is the x for those taking the
preparatory course.
2. The mean of our frequency distribution (the sampling distribu-
tion), X, according to the central limit theorem, equals yw for all the
ninth-graders.

3. The standard deviation of our frequency distribution, which for a


sampling distribution is called the standard error, according to the
central limit theorem equals o for all the ninth-graders divided by the
square root of our sample size: o//7.

Substituting this information from the central limit theorem for the
words in our equation, we get

ms (1) Xcourse — (2) [al]

(3) [oan //7]


Simplifying,

_ Xcourse — Hall
Oa //n

Or, in general terms,

Cet
o//n

Thus, the formula for the one-sample z test of significance is really


a recasting of the basic formula z = (x — 44)/o to apply to the sampling
distribution. The central limit theorem enables us to find a z value on the
sampling distribution from data pertaining to the population and the sample.
This is illustrated in Figure 8.9, which shows our sampling distribution (not
drawn to actual scale) and its components.
We now see why a directional H, is dubbed a one-tailed H,—we only
make use of one tail on the sampling distribution. Without a directionality
assumption, we would move out from the mean of the sampling distribution
Probability Distributions and One-Sample z and t Tests » 243

Figure 8.9 = The Sampling Distribution of Sample Means for Samples 7 = 100
Based on Competency Exam Data (Hypothetical)

eee ie
Cx
A/S nex 100” weAD: ®

Since z=1.5 < 1.65


<— Tail area >.05
Cannot reject Hg

65 6667 6869 ig 72 {7475

X =P ay=70 x course = 73

toward both the left and the right, examine the size of both of the tails by
comparing the absolute value of z, obtained to Z,critical? and pay the price of need-
ing a larger Zjraineq than is needed when using only one tail.

Review

Before we proceed, let us review the fact that in using the central
limit theorem, we are working with three separate frequency distributions:
the population, the sample, and the sampling distribution of sample
means. We are given information about the first two distributions. The cen-
tral limit theorem then enables us to take data from those two distributions
and make use of the properties of the sampling distribution. We know the
following:

1. The frequency distribution for variable x for some population. We


assume that this distribution is normal. We know its mean yp and its
standard deviation o.

2. The frequency distribution of a particular random sample that we


have drawn. We know its size 7 and its mean xX. The variable x in our
sample is the same variable x in our population.
244 << STATISTICS FOR THE SOCIAL SCIENCES

3. The sampling distribution of sample means. (You never see this


distribution; you just make use of it!) There exists a separate sam-
pling distribution of sample means for each possible sample size
(each n). For any given 7, this represents the frequency distribu-
tion of all possible sample means from all possible samples drawn
randomly from that population whose mean and standard deviation
along variable x are yw and o, respectively.

For the specific sample that we have drawn, our sample mean x will be
one point (one value ofX) on that sampling distribution. The central limit
theorem enables us to find the distance from the sample’s mean to the
population’s mean, expressed as standard errors or standard deviations of
the sampling distribution. Since the sampling distribution is a normal curve,
we may determine the probability of our sample’s x reflecting a population
whose mean is w and, based on that probability, either retain or reject our
null hypothesis.

THE NORMALITY ASSUMPTION

Note that the central limit theorem assumes that the population we are
studying is normally distributed along variable x. This is called the normal-
ity assumption. If it is true, the sampling distribution of sample means will
be a normal distribution, and we may make use of the z formula to test
for statistical significance. (Note that nothing requires that our sample be
normally distributed.) What if we know that the population is not normally
distributed, or more realistically, what if we have no basis for making a
normality assumption about the population in the first place? Even in such
cases, if our sample’s size is large enough, we may still be able to make use
of the central limit theorem due to the law of large numbers.
UIT NCES ORIEN INERT LUBE
INLINE SES ANLEEOTE
OLS EIT

Normality assumption The assumption that that the population being studied is
normally distributed along variable x.
nan CRA

The law of large numbers states that if the size of the sample, 7,
is sufficiently large (no less than 30; preferably no less than 50), then the cen-
tral limit theorem will apply even if the population is not normally distributed
along variablex. Thus, if 7 is large enough, the population distribution need
not be normal and could, in fact, be anything: skewed, bimodal, trimodal,
anything. When 7 is large enough, we relax the normality assumption for
our population, but the sampling distribution of sample means will still be a
normal curve, and the central limit theorem will still apply.
Probability Distributions and One-Sample z and t Tests 245

Law of large numbers A law that states that if the size of the sample, n, is
sufficiently large (no less than 30; preferably no less than 50), then the central limit
theorem will apply even if the population is not normally distributed along variable x.

How large must 77 be to relax the normality assumption? The figures given
in the above theorem are rather arbitrary; other sources give other cutoffs. In
fact, in some texts of statistics for psychology (which often only requires small
samples or small experimental and control groups), the minimum sample size
is as low as 15, but that is probably too low. Perhaps we ought to put it this way:

UF Then:
n 2 100 It is always safe to relax the normality assumption.
50 <n < 100 It is almost always safe.
50'S 7 < 50 It is probably safe.
YoU It is probably not safe.

In social science survey research, our sample sizes are generally large
enough to make use of the law of large numbers. This is particularly fortu-
nate, since in actual research all too often, the issue of the normality
assumption is not adequately addressed.
Let us look at an example. At a small liberal arts college, an index of
support for civil liberties, ranging from 0 (least supportive) to 10 (most
supportive), was pilot tested on the entire student body, yielding a mean of
7.5 and a standard deviation of 1.5. A random sample of 100 students who
had been the direct victims or close relatives of victims of serious crimes was
also given the test, and their mean score was 7.2. May we conclude that for
all similar victims, the support score for civil liberties differs in general from
the population of all students at that college?
Our hypotheses are

AL: May = Kyictims

Hy: bay % victims (NO directionality assumed)

We
Since 2 =100, we may relax the normality assumption for the population.
have all necessary data for a one-samp le Z test.

Cote aa Xvictims ~ Hall is = ID


oa

o/J/n Oa // victims 1.5// 100

ee 100
= —2.00
Sars de ks
246 << STATISTICS FOR THE SOCIAL SCIENCES

We compare the absolute value ofz to the two-tailed Z,,,,,.., values:


| eis OG reject H,
200-2258, Pees

We conclude, therefore, that the civil liberties support score for all seri-
ous crime victims at this college is lower than the average for the college as
a whole (p < .05). (The sampling distribution is shown in Figure 8.10.)

THE ONE-SAMPLE ¢ TEST

We know that to do the one-sample z test, we need to know or be able


to hypothesize two population parameters, 4 and 6. What could we do in
the unlikely event that we know yu but not 6? Initially, the sample standard
deviation s was assumed to be a good estimate of 6, so s was substituted
when o was unknown. Once the sample mean X had been calculated, s was
generated using the definitional formula

or one of several possible computational formulas.

Figure 8.10 The Sampling Distribution of Sample Means for Civil Liberties
Support Scores

lz eriticat 105 level|


i = 1.96 <|z|=2.0
A 7~ | 27> 1 Thus p, the total tail area, < .50.
1.96] 1.96
Thus, reject Ho.

Ieee 15 1.5
Re)
ae -h0) em

ll
Probability Distributions and One-Sample z and t Tests » 247

However, it was discovered that, particularly when the sample size 7 was
small, calculating z with s produced inaccurate conclusions. A British quality
control expert* working for a Dublin brewery discovered that by calculating
a different estimate of 6 from sample data, a better test of significance could
be developed. This new best “unbiased” estimate of 6, which we designate
6 (read as “sigma-hat,” because sigma is wearing a hat), is created when we
substitute 7 — 1 for m in the standard deviation formula.

Sigma-hat (6) An estimate of sigma.

lob II

This new test of significance is called the ¢t test to differentiate it from the
z test; note that the formulas are the same except that 6 is substituted for o

xX —
~— 6fln
When 7 is large, the substitution of 6 for s makes very little difference, but
as m gets smaller, 6 and s diverge, causing a likewise divergence between f
(using 6) and z (using s to estimate 0).

ttest A test of significance similar to the z test but used when the population’s
standard deviation is unknown.

The sampling distributions of t and z also differ. In the case of the z test,
the sampling distribution of sample means is a normal curve. Since the value
of each sample mean can be expressed as a z score (indicating the distance
x is from yw in terms of standard errors), the sampling distribution of sample
means is the same as the distribution of all the z scores from all the theo-
retically possible sample means that make up the sampling distribution.
Thus, the sampling distribution of z (all the zs from those sample means) is
also a normal curve.
If we take the same means in our sampling distribution and calculate ¢
scores instead, the sampling distribution of ¢ (all the ¢s from those sample
means) is a normal distribution only when the sample sizes are above 120.
As the sample sizes fall below 120 (give or take), the sampling distribution
begins to be flatter than a normal curve (say platykurtic, if you want to
impress your friends). When the curve is flatter than a normal curve at its
peak, the tails are also larger than those of a normal curve. (The effect is sim-
ilar to pushing a balloon down from its top, thus displacing the air to the
sides as we press.) As ” gets smaller, the peak of the sampling distribution
248 << STATISTICS FOR THE SOCIAL SCIENCES

gets flatter, and its tails get larger. The important consequence is that as 7
gets smaller, we must go ever-greater distances away from the mean to get
a tail area equal to .05 proportion of the area under the curve.
Figure 8.11 shows the changes in the critical values of ¢ (.05 level,
one-tailed) as m decreases. At 7 = 121, the sampling distribution is nearly
a normal curve, and ¢critical is 1.658, only slightly larger than Z,,,,,..; (05 level,

Figure 8.11 Changes in the Sampling Distribution of tas Sample


Size Decreases

N=21
ai=20
Somewhat
flatter than a
normal curve

|
|
|
|
|
|
|
|
|
Much |
|
flatter than a
|
normal curve | p=.05
|

# —> X=
t=2.015
Probability Distributions and One-Sample z and t Tests j» 249

one-tailed), which is 1.65 (actually 1.645 before rounding). In fact, as 7


increases above 121, the critical values of z and ft get ever closer to each
other. As 7 gets extremely large, approaching infinity as a limit, the critical
values of z and t become the same. However, as 7 drops below 121, the
iticg Value gets larger. In other words, we have to go farther out to get a tail
l
with .05 of the area under the curve in it. By the time 7 = 21, tcritical has gone
from 1.658 to 1.725, and at 7 = 6, t.,..,., has risen to 2.015.
In comparing the critical values of ¢ to those of z, bear in mind that the
sampling distribution of z is always a normal distribution, and its critical
values remain constant, independent of sample size. By contrast, the critical
values of ¢ depend on sample size. At best, when 7 is large, the critical values
of t are almost as small as those of z. But as 7 decreases, the critical values of
t get larger, making it harder to reject the null hypothesis. Thus, if we know
o and can therefore do a one-sample z test, we always do the z test, not the
t test. We do the one-sample ¢ test only if o is unknown and we must estimate
it with G. In fact, when 7 gets large (say 30 or more), many statisticians advo-
cate the use of the z test, with s substituting for o in the formula, rather than
the use of the ¢ test. But with a smaller 7 where o is unknown, we must
always do the ¢ test and retain the normality assumption for the population.

DEGREES OF FREEDOM

Note that in Figure 8.11 under each of the three reported ms—121, 21, and
6—is another number labeled df which is one less than m—120, 20, 5. As we
learned in Chapter 7, the df stands for degrees of freedom, a number we
generate to make use of a table of critical ¢ values. In the case of the one-
sample ¢ test,

df=n—1

Degrees of freedom A number that is generated to make use of a table of critical


values.
EN
EBCEESSISNSOSEEE

We need to find the degrees of freedom in order to find the critical values
of t against which we compare our obtained ¢. As noted, the sampling distrib-
ution of t changes from a normal curve as 7 decreases, and thus the critical
values change as well. As we see in Figure 8.11, at 120 degrees of freedom (77 =
121) we need at of 1.658 to have one tail on the sampling distribution with a
05 area. By the time degrees of freedom drops to 5, we need at of 2.015.
Tables of critical values for all tests of significance beyond the z test
require that we first calculate a degrees-of-freedom figure to make use
250 € STATISTICS FOR THE SOCIAL SCIENCES

of the table. Why find df? Why not base the tables on 7 as we did the
sampling distributions in Figure 8.11? The simplest answer to the question
is that there are several formulas that generate ¢ scores, not just the one
presented in this chapter. Likewise, for each of the different ¢ formulas,
there is a separate degrees-of-freedom formula. The formula df=n — 1 is
used only for the one-sample ¢ test presented here. In the next chapter, we
will discuss some of the other ¢ formulas, each having its own degrees-
of-freedom formula, but all making use of a common table of critical values
of t. Without degrees of freedom, we would need a separate table of critical
values for each separate formula.
There is a mathematical meaning to the concept of degrees of freedom,
having to do with how many numbers are free to vary in a formula. For
instance, if x,+.%, +x, = 10 and you let any two of the scores vary (say we
make x, = 2 and x, =5), then the remaining value ofx is fixed. Since 2+5=
7 and 7 + x, = 10, once x, and x, are determined, x, can take on only one
value. In this case, x, = 3. So three unknowns adding up to a fixed sum has
two degrees of freedom. Only two of the unknowns are free to vary. At the
level of applied statistics that we cover in this book, it is not really necessary
to know the definition of degrees of freedom to make use of the concept.
So we will simply move on, referring the curious to more advanced texts.
For our purposes, degrees of freedom are simply numbers that we must
calculate to make use of critical values tables for ¢ and the other tests of
significance to be encountered later.

THE ¢ TABLE

The table of critical values of 4,found in Table 8.2 and also in the Appendix,
is simple to use. At the top are levels of significance for a one-tailed test (a
directional H,), and below it are the corresponding levels for a two-tailed
test. Thus, ¢oitica, ONe-tailed at the .10 level is the same as f.,,,.., two-tailed
at the .20 level. The one-tailed probability levels are always one half of the
corresponding two-tailed levels.
Since we always begin by comparing fjrainea tO Loriticas at the .05 level, we
first isolate the appropriate .05 column for whicheverH, (one-tailed or two-
tailed) we are using. Then we go down the df column on the far left until we
come to the number that we found in the df formula. Noting the values
highlighted earlier in Figure 8.11, if df is 120, we go all the way down the
df column until we find 120. We then move across the row until we are
under the .05 level for a one-tailed test. At the intersection of the 120 row
and the .05 column, we find the critical value of ¢ 1.658. Likewise, in the
same .05 column, we find the ¢(iticq, Of 1.725 in the row for 20 degrees of
freedom and 2.015 in the row for 5 degrees of freedom.
Probability Distributions and One-Sample z and t Tests » 251

Table 8.2 Distribution of ¢

Level of significance for one-tailed test

10 105) 025 O01 005 0005

Level of significance for two-tailed test

df 20 10 05 02 01 001
1 3.078 6.314 12.706 31.821 63.657 636.619
2 1.886 2.920 4.303 6.965 9.925 31.598
3 1.638 2.353 3.182 4.541 5.841 12.941
4 1.533 2.132 2.776 3.747 4.604 8.610
5 1.476 2.015 2.571 3,365 4.032 6.859
6 1.440 1.943 2.447 3.143 3.707 5.959
7 1.415 1.895 2.365 2.998 3.499 5.405
8 1.397 1.860 2.306 2.896 3.355 5.041
9 1.383 1.833 2.262 2.821 3.250 4.781
10 1372 1.812 2.228 2.764 3.169 4.587
gh 1.363 1.796 2201 2.718 3.106 4.437
12 1.356 1.782 2.179 2.681 3.055 4.318
13 1.350 {77 2.160 2.650 3.012 4.221
14 1.345 1.761 2.145 2.624 2.977 4.140
15 1.341 1.753 2.131 2.602 2.947 4.073
16 1.337 1.746 2120 2.583 2.921 4.015
17 1.333 1.740 2.110 2.567 2.898 3.965
18 1.330 1.734 2.101 2.552 2.878 3.922
19 1.328 1.729 2.093 2.539 2.861 3.883
20 1.325 1.725 2.086 2.528 2.845 3.850
on 1.323 L721 2.080 2.518 2.831 3.819
20 1.321 L717 2.074 2.508 2.819 3.792
23 1.319 1.714 2.069 2.500 2.807 3.767
24 1.318 17d 2.064 2.492 2797 3.745
25 1.316 1.708 2.060 2.485 2.787 3.725
26 1.315 1.706 2.056 2.479 2.779 3.707
27 1.314 1.703 2.052 2.473 277i 3.690
28 1.313 1.701 2.048 2.467 2.763 3.674
29 1.311 1.699 2.045 2.462 2.756 3.659
30 1.310 1.697 2.042 2.457 2.750 3.646
40 1.303 1.684 2.021 2.423 2.704 3.551
60 1.296 1.671 2.000 2.390 2.660 3,460
120 1.289 1.658 1.980 2.358 2.617 3.373
co 1.282 1.645 1.960 2.326 2.576 3.291

SOURCE: Abridged from R. A. Fisher and F. Yates, Statistical Tables for Biological,
Agricultural and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint
of Pearson Education.
252 << STATISTICS FOR THE SOCIAL SCIENCES

Under the df= 120 row, we note the symbol for infinity (an eight that
has gone down for the count). In this case, “infinity” is any df above 120.
Here, the sampling distribution has become (or is in the act of becoming) a
more perfect normal curve. Note that at this point, there is no difference
between the critical values of ¢ and those of z.
If you cannot find the df that you need in the table, go to the nearest
critical value that makes it harder to reject H). In the case of Table 8.2, move
up to the next lower df For instance, if the dfis 35, a number not presented
in the table, go up to 30 df and use those critical values. Thus, if the ¢
obtained in a one-tailed test at 35 df were 1.7, you would compare it to the
.05 critical value at 30 degrees of freedom, 1.697. Since 1.7 is greater than
1.697, you would reject H,. What if the obtained ¢ were 1.690? That would be
less (barely) than 1.697, and you could not reject H, using this table.
However, you would be right in assuming that had you known f,,,.,... at 35
degrees of freedom, there would be a good chance that it would be equal to
or less than your ¢t of 1.690. In this case, consult a book of tables for statisti-
cians, which would have a more complete ¢ table than the one used here.?
If the obtained ¢ exceeds f,,,,,.., at the .05 level, you then compare it to
the critical values to the right of the .05 column. Following the same proce-
dure used for the z test, you make your probability statement by seeing how
many critical values are less than the obtained ¢. The only difference is that
in the ¢ table, there are critical values for levels other than .05, .01, and .001.
Suppose at 60 df we obtain a ¢ value of 3.0 using a nondirectional H,. Going
down the .05 level column, for the two-tailed test, we see at the 60 df row a
critical value of 2.000. We can reject H,. We then compare our 3.0 obtained
¢ to the critical values to the right of the 2.000 we exceeded. We exceed the
2.390 (.02 level) and the 2.660 (.01 level) but not the 3.460 critical value at
the .001 level. Thus, we report p < .01. Had this been a one-tailed test, we
would be reportingp < .005.

AN ALTERNATIVE ¢t FORMULA

We have been using the following formulas:

Ree iid tral


Me: Rae
where
Probability Distributions and One-Sample z and t Tests 253

Suppose you did not have access to but did know the original sample
standard deviation of

Rather than recalculating, you may make use of the s in a modified ¢ formula:

x — 4
t = ——— _ df=n-1
s//n—1 poset

Again, remember that if you get the standard deviation from either a com-
puter printout or a calculator with a standard deviation function built in,
consult the appropriate manual to find out how that standard deviation
was calculated to determine whether you have an s or a G. Then pick the
appropriate ¢ formula to use.

A z TEST FOR PROPORTIONS

The formula for the z test for sample means may be modified to test the
difference in proportions in a sample compared to the equivalent difference
in proportions in a population.
SLSR OE IE EDL EC EEDA SOY LT ESOL

z test for proportions A z test designed to test whether the difference between
proportions in a sample reflects the difference in the population.
panei aH IE
TS OE OOS TEES

For instance, suppose that in some small community, the proportions


of minorities (people of African or Hispanic origin) make up 20% (.20 pro-
portion) of the population. The new school superintendent suspects that
minorities are underrepresented among the 100 teachers in her public school
system since there are only 15 minority faculty, 15% or a .15 proportion. For
such a problem,

Pa Pp

VP pQp/n

where

P, =the proportion of minorities in the sample = .15,


P,, = the proportion of minorities in the population = .20,
Q, = the proportion of nonminorities in the population = 1—P, = 1—.20
= OU,
n =the size of the sample or group being studied = 100.
254 @ STATISTICS FOR THE SOCIAL SCIENCES

ene

Ay: - =i

1 OEM ee oe

i =
JPpQpin —J©20)(.80)/100
=,05 = =.U5
= = = —— = -1.25
J .16/100 ~/.0016 .04
Using the [Link] At the .05 level of 1.65, we cannot reject H, since
1.25 < 1.65. We cannot conclude that minorities are underrepresented
among the teachers.

INTERVAL ESTIMATION

We have already discussed the fact that if we did not know o, our best
estimate of it from sample data would be 6. Likewise, our best estimate
of w would be X. Suppose we wanted to estimate yz from x. We know from
the sampling distribution of sample means that not all sample means will be
exactly equal to w, even though our one x is the best estimate of that para-
meter. With interval estimation, we establish an interval of scores called a
confidence interval, and we state with a certain level of confidence that
the w will fall within the limits of the interval we created.

Interval estimation An interval of scores that is established, within which a


population’s mean (or another parameter) is likely to fall, when that parameter is
being estimated from sample data.

Confidence interval (for means and proportions) An estimated interval within


which we are “confident”—based on sampling theory—that the parameter we are
trying to estimate will fall.

For instance, we can see from our sampling distribution that with no
directionality assumption, 95% of all sample means lie between w and £1.96
standard errors. Likewise, 99% lie between jz and 2.58 standard errors. The
number of standard errors corresponds to the two-tailed Z,,,, at the .05
and .01 levels, respectively. Also, 99.9% of all sample means lie between
and +3.29 standard errors, and 3.29 is the critical z at the .001 level. Suppose
we would be satisfied to find the interval within which 95% of all sample
means would fall. We build an interval around the X and assume that p will
Probability Distributions and One-Sample z and t Tests 255

fall within that interval. We call this the 95% confidence interval, our level
of confidence corresponding to the percentage of all means falling within
the interval. Thus, we are 95% confident that w will lie in the interval
between x — 1.96 0, and x + 1.96 o,.
Remembering that we already know that o-=0/\n, we find our
confidence interval by the following formula:

X + 1.96(0//n)
Suppose X = 55, 6 = 10, and m = 64. The upper limit of our interval would be

X + 1.96(0/V/n) = 55 + 1.96(10/V64)
= 55 + 1.96(10/8)
= 5542.45

= 57.45

Our lower limit would be

X — 1.96(/V/n) = 55 — 1.96(10/V'64)
= 55 — 1.96(10/8)
= 55 — 2.45
= 52.55

Thus, the 95% confidence interval for estimating w is 52.55 to 57.45.


We know that 95% of all sample means fall within the interval, so we are
95% confident that p will be between 52.55 and 57.45.
Suppose we wanted a greater level of confidence, 99%. The price we
would pay for it would be a wider confidence interval. For our upper limit,

X + 2.58(0//n) = 55 + 2.58(10/V64)
= 55 + 2.58(10/8)
= 55 + 3.23
= 58.23
and for our lower limit

X — 2.58(0//n) = 55 — 2.58(10/V64)
= 55 — 3.23
is ol
a

We are 99% confident that p falls between 51.77 and 58.23.


256 STATISTICS FOR THE SOCIAL SCIENCES

Ifois unknown, which is generally the case, we may do exactly the same
procedure with the ¢ test using either of the following formulas:

pL =X + tcritical (6//n)

or

=X Leritical (S/V 7 — 1)

The nondirectional ¢.,;..4) at df=n — 1 at the .05 level would be used for a
95% confidence interval, the tosis , at the .01 level would be used for a 99%
confidence interval, and so on.

CONFIDENCE INTERVALS FOR PROPORTIONS

Imagine that you are a campaign manager of a presidential candidate in


a two-person race. A telephone survey of 900 voters gives your candidate
a 53% lead over the opponent. How likely does that percentage lead
reflect the electorate? You seek to construct a 95% confidence interval
around the .53 proportion that your candidate received in the sample.
The formula we use is

Ps £1.96,/PpOp/n

The 1.96 is the appropriate critical value of z—in this case, at the .05
level since we chose a 95% confidence interval.

P. = your candidate’s proportion of support in the sample = .53.

P, = your candidate's proportion of support in the population, which


we estimate with P, Thus, P, =P, = .53.

Q,= the opponent’s proportion of support in the population, which


we estimate from the sample by subtracting P. from 1. Thus Q,=
PHP, =a PS 155 Say
nm = the number of cases, which must equal or exceed 5/min(P, LP),
that is, 5 divided by whichever is smaller, P, or 1 — P.

Thus, the upper limit of our 95% confidence interval is


Probability Distributions and One-Sample z and t Tests » Psy

Ps + 1.96,/PpQp/m = .53 + 1.96y/(.53)(.47)/900


53 + 1.96,/.2491/900
= 53 + 1.96V.00028
53 + 1.96(.0167)
53 + .0327
II 53 + .03
= 1.56

The lower limit would be

P, — 1.96,/P5Op/n = .53 — 1.96y/(.53)(.47)/900


— Do nD
we) )

So our confidence interval ranges from .50 to .56. Since in percentages,


this is 50% to 56%, a range of 6 percentage points, we report that according
to our poll, our candidate has a 53% lead, but our margin of error is plus or
minus 3 percentage points. Our candidate could receive as little as 50% or
as much as 56%. If we had chosen a 99% confidence interval, our confidence
interval would be larger and so would the margin of error reported.
When 7 is small or if we want to be particularly sure of our estimate, it
is safer to make a more conservative estimation of P,, than to use P. Here,
we assume that each candidate has half of the vote. Thus, P, = .50 and Q, =
50. This will yield a larger confidence interval than any other estimate of P,
would generate. By widening the interval, we minimize the risk in making
our estimate. Suppose P, were .53, but 7 = 150 instead of 900. We estimate
P, and Q,, as .50, respectively. For the upper limit,

P, + 1.96,/PpOp/n = .53 + 1.96,/(.50)(.50)/150


= 53 + 1.96,/.25/150
= 53 + 1.96/.00167
— 53 + 1.96(.041)
— Oo eue
= 61
And our lower limit would be

53 =..08 = .45

Here, our 95% confidence interval ranges from .45 to .61, and we have an 8
percentage point margin of error.
258 & STATISTICS FOR THE SOCIAL SCIENCES

MORE ON PROBABILITY

Suppose we have developed a scale to be used in a survey. This scale mea-


sures the extent to which the respondent is aware of and knowledgeable
about HIV and AIDS. Assume that the scale ranges from a low of 0 to a high
of 100 and is a normal distribution with a mean of 50 and standard deviation
of 15. We may thus apply the z formula to this distribution to determine the
proportion of cases falling within a specified range of scores. Let us use this
scale to extend our discussion of probability.
We begin by outlining some new notations and defining them.

P(A) =the probability of outcome A occurring.

P(A or B) = the probability of either outcome A or outcome B occurring.

P(A and B) = the probability of both outcomes A and B occurring jointly.

P(A |B) =the probability of outcome A occurring given that outcome B


has already occurred (conditional probability).

Conditional probability —The probability of outcome A occurring given that


outcome B has already occurred.

Let us illustrate using our AIDS awareness scale. Suppose Outcome A is


the probability of selecting an individual with an AIDS awareness score of 70
or above. Since x = 70, wu= 50, and o = 15, we apply the z formula and find
z= 1.33. Looking at Table 8.1, column C, we find a probability of .0918. Thus,
P(A) = 0918.
Let outcome B be the probability of selecting someone with an AIDS
awareness score of 40 or below. Plugging into the z formula, we obtain a z
of —0.66 and find a probability of .2546 from Table 8.1. Thus, P(B) = .2546.

The Addition Rule

Suppose we would like to know the probability of selecting someone


whose AIDS awareness score is either 70 or above or 40 or below, P(A or B).
Now outcomes A and & are known as mutually exclusive outcomes. If one
has an AIDS awareness score above 70, one cannot also have an AIDS aware-
ness score below 40, When outcomes are mutually exclusive, a rule known
as the addition rule tells us that

P(A or B) = P(A) + P(B)


Probability Distributions and One-Sample z and t Tests 259

In this case,

P(A or B) = P(A) + P(B) = .0918 + .2546 = 3464

Addition rule A rule by which when outcomes are mutually exclusive, the
probability of either outcome occurring is the sum of the probabilities of each
outcome occurring.

If we had included a third outcome, outcome C, such as awareness


between 50 and 55, we would calculate a z of 0.33 and, looking this time at
column B of Table 8.1, find a probability of .1293.

PUG) = 21295

Therefore,

P(A orB or C) = P(A) + P(B) + P(C)


= 0918 + .2546 + .1293
= A757,

When our events are not mutually exclusive but overlap, we must apply
a more complex addition rule. Suppose outcome A remains a score of 70
and above, and we add another outcome, outcome D. If outcome D is the
probability of selecting a respondent with AIDS awareness between 50 and
75, z will be 1.66, and column B of Table 8.1 will yield a probability of .4515.
This time, however, we cannot simply add P(A) to P(D) to find P(A or D)
since our outcomes are no longer mutually exclusive. Anyone with a score
between 70 and 75 will belong jointly to both outcomes. To account for this,
we must expand the addition rule as follows:

P(A or D) = P(A) + P(D) - P(A and D)

We will see ina moment how P(A and D) is determined, but for now assume
that we are told that it is .0414. Therefore,

P(A or D) = P(A) + P(D) - P(A and D)


0918 + .4515 — .0414
5019

Note that in the first example, P(A or B), A and B had no overlap, so
P(A and B) = 0. Applying the longer addition rule,
260 STATISTICS FOR THE SOCIAL SCIENCES

P(A or B) = P(A) + P(B) — P(A and B)


0918 + .2546 —0
3464

This matches the result found earlier.

The Multiplication Rule

As with the addition rule, there are two forms of the multiplication
rule, the rule that we use to find P(A and D). The simple form of this rule
applies when the outcomes or events are independent of one another—
when neither event influences the probability of the other event occurring.
Symbolically,

P(A |D) =P(A) and PO Ay= PD)

Multiplication rule A rule that is used to find P(A and D).

In these two events, A and D are independent; that is, neither event will
affect the probability of the other event’s occurrence. In our example, deter-
mining the probability of selecting someone whose AIDS awareness is 70 or
more has no impact on determining the probability of selecting someone
with an awareness score between 50 and 75. Two z scores are calculated
independently of one another.
In the case of independent events, the multiplication rule becomes

P(A and B) = P(A) x P(B)


In our example,

P(A and D) = P(A) X P(D) = (.0918)(.4515) = .0414

That was how the value of P(A and D) used in the addition rule above was
determined. Like the addition rule, the multiplication rule can be extended
to more than two independent events.

P(A and D and £) = P(A) x PD) x P&)

What about nonindependent events? Let us assume that anyone with a score
of 65 or greater has high AIDS awareness. Here z = 1.00, and column C of
Table 8.1 shows that the probability of selecting a high-awareness person is
.1587. If we have a finite group of 13 individuals, we would expect to find
.1587 x 13 or 2.06 high-awareness scores.
Assume, therefore, that we have 13 people, 2 of whom have high
AIDS awareness. What is the probability of making two selections from the
Probability Distributions and One-Sample z and t Tests » 261

group and selecting the 2 high-awareness individuals? The events are


nonindependent since the outcome of the first selection has an impact on
the second selection. The probability of getting a high-awareness scorer on
the first draw would be 2/13 or .1538.
If we select a high scorer on the first draw, the probabilities change on
the second draw: There are now 12 people, 1 of whom is a high-awareness
scorer. The probability of selecting a high scorer on the second draw after
having selected a high scorer on the first draw, P(H, |H,), is 1/12 or .0833.
(Here H stands for high scorer; 1 and 2 for the first and second draws,
respectively.) So

P(A, and H,) = P(A,) x PA, |A)


= (.1538)(.0833)
= .0128

In more general notation,

P(A and B) = P(A) x P(B |A)

or

P(A and B) = P(B) x P(A |B)

If two events are statistically independent events, then

P(A |B) = P(A) and PB |A) = P(B)

Thus, we return to the simpler formula for the multiplication rule.

P(A and B) = P(A) x P(B |A) = P(A) x PB)

PERMUTATIONS AND COMBINATIONS

Earlier in this chapter, we discussed how many different samples of 7 = 3 we


could draw from a population where 7 = 5. Recall that there were 60 possi-
ble samples when the order of selection was considered. When we drew
first a5 then a 4 then a3, we considered it a different sample than when we
drew first the 4 then the 5 then the 3. The total possible samples that can
be drawn from a population when the order of selection is a factor is called
a permutation. The total possible samples when the order of selection is
ignored is called a combination. In the earlier example, there were 10
different possible samples when the order of selection was ignored.
262 @ STATISTICS FOR THE SOCIAL SCIENCES

Permutation The total possible samples that can be drawn from a population when
the order of selection is a factor.

Combination The total possible samples when the order of selection is ignored.

The formulas for finding permutations and combinations are presented


below, whereN is the number of items in the population, K is the size of the
sample to be drawn, and m! (read n factorial) is a number times each
number lower than itself down to 1. (For example, 5!=5x4x3x2x1=
120. 4 =4 x 3oc2 < b=24-51=3%2x« l= Gand soon, Note that wenever
include zero in the multiplication or our answer would always be zero!)
For permutations, indicated by the letter P

N!
2 yess Daas
A NR

In our example,
NV=5 and Ki=3,'so

pre N! 2 5! 2 ee ee
TNS "Gea ae oe oe ee
For combinations, indicated by the letter C,

Ch eee
RLQY =)
In our example,

Nes N! = 5! +f 5! DG Se
= RN=k B6—o Sel Gaerne op
26 _ 120 46
SKOCr™ aa
CONCLUSION
In Chapters 7 and 8, all the basic elements of tests of statistical significance have
been presented in a time-honored sequence, moving from the normal distri-
bution in its basic form to the one-sample z test and then to the one-sample
t test. As stated earlier, every test that follows in this text also follows the same
logical assumptions and basic procedures, starting with the formulation of A,
and H,, calculating the df (if appropriate), comparing the obtained value to crit-
ical values of that statistic, reaching a decision as to whether or not to FejCcuia
and, if H, is rejected, formulating the appropriate probability statement.
Probability Distributions and One-Sample z and t Tests » 263

However, the tests presented so far have only limited value in that since
they are one-sample tests, we are comparing data from that one sample to
data from a population. Rarely do we know population parameters such as
u and o, although it might be possible to estimate them, and rarely do we
know if it is valid to assume that these populations, in fact, are normally dis-
tributed for the variable in question. More often, we are comparing the
means of two or more samples, and we know no population parameters at
all. Often, we have problems involving nominal or ordinal levels of mea-
surement when a comparison of means is inappropriate. We cover tests for
these purposes in the following chapters.

Chapter 8: Summary of Major Formulas

The z Formula for the Area Under a Standard Normal Curve

Sigma-Hat (estimate of a population standard deviation from sample data)

Cee
6 =
n—1

r The ¢ Test of Statistical Significance (calculated with sigma-hat)

cy!) N= ee A
= Shin
The ¢ Test of Statistical Significance (calculated with s)

(es ian il

The z Test for Difference of Proportions

Pua

VPpQp/n
264 4 STATISTICS FOR THE SOCIAL SCIENCES

EXERCISES
Exercise 8.1
An index of cognitive awareness is normally distributed with a mean of 1 = 8.9 and
a standard deviation of 6 = 3.1.

1. What proportion of people would be expected to have awareness scores of 14


and above?
What proportion would have scores between 8.9 and 14?
What proportion would have scores below 14?
What proportion would have scores of 8.5 and below?
What proportion would have scores between the mean and 8.5?
What proportion would have scores ranging from 8.5 to 14?
What proportion would have scores either less than 8.5 or greater than 14?
DAR
WN
SNRemembering that the areas under the normal curve are also the probabilities
of randomly selecting someone with a particular characteristic, what is the
probability of randomly selecting a person with an awareness score of 12.5 or
more?
9. What is the probability of randomly selecting a person with an awareness
score between 6 and 8.9?
10. What is the probability of randomly selecting someone whose awareness level
is between 6 and 12.5?

Exercise 8.2
Each of the following problems requires either a one-sample z or a one-sample
t test. Select the appropriate test and perform it. Assume a nondirectional H, unless
the wording of the problem suggests otherwise. For each test, indicate whether or
not the normality assumption may be relaxed for the population. In doing the ¢ test,
make sure you are using the appropriate formula; that is, are you given Gor s?

1. Suppose you know that for the entire United States, the mean age of the popu-
lation is 32, with a standard deviation of 14.5 years. Since many retired people
move to Florida, you believe that the mean age of all Florida residents is greater
than that for the United States as a whole. You randomly select a sample of 144
Floridians and obtain a mean sample age of 34.
2. For the Miami metropolitan area, the mean age of a random sample of 25 resi-
dents is 36.5, with a standard deviation of s = 16 years. Compared to the United
States (data given in Part 1), what may we conclude about Miami residents?
3. Suppose the sample size in Part 2 had been n = 64. What would your conclu-
sion be?
Probability Distributions and One-Sample z and t Tests jp 265

4. A scale designed to measure support for gun control legislation has been
developed. It ranges from 0 to 10, with 10 meaning strongest support for such
actions as outlawing “Saturday night specials” and semi-automatic weapons.
Suppose it has been determined that for the entire population of the state of
Maryland, the mean support score is 6.0. A random sample of 100 residents
of Maryland’s Eastern Shore yields a sample support score of 4.8 with 6 = 4.0.
What do you conclude?
5. For Baltimore County, a random sample of n = 81 has a mean of 7.0 and a
standard deviation of 6 = 4.5. (For this problem and the ones that follow, use
the population figures given in Part 4.) What do you conclude for each one?
6. For Baltimore City, a random sample of 31 residents produces a mean of 8.0
and a standard deviation of s = 3.0.
7. Arandom sample of 170 members of the National Rifle Association who live
in Maryland yields a mean of 1.5 and an s = 1.25.
8. To ascertain the attitudes of all residents of the city of Cumberland, Maryland,
a random sample of 9 members of that city’s police department was inter-
viewed. The sample’s mean was 8.5, and its 6 was 3.0.

Exercise 8.3
At a state’s maximum-security penitentiary, all inmates have taken a battery of
psychological tests. Following are the means and standard deviations for several
selected indices developed from those tests.

Index p= (Gie

VIO Attitudes supporting violence d7.1 15.2


ISO Feelings of isolation from others 56.2 32.4
“RAC Attitudes of tolerance toward other races 62.5 31.6

A random sample of 50 inmates at low- and medium-security institutions in the


same state yields the following:
a

Index xe
VIO 3/4
ISO 50.7
RAC 713

Using one-sample z tests (two-tailed), test for significant differences between these
two groups for
1. ViO
2, 180
3: RAC
What are your conclusions?
266 € STATISTICS FOR THE SOCIAL SCIENCES

Exercise 8.4
Following are the maximum-security penitentiary population means for three other
indices.

Index p=
REM Remorse for the victim of the crime 74.9
DET Determination to commit no further crimes 48.6
ALI Alienation from societal norms 42.5

~ For the low- and medium-security sample, n = 50, the statistics are as follows:
Index X= o=
REM 78.2 16.2
DET 60.2 33:0
ALI 36.6 35.)

Using a nondirectional one-sample t test, test for significance and state your
conclusions.
1. REM
2. DET
3. All

Exercise 8.5
Suppose that for the population, it is known that 51% are women and 49% are men.
Suppose random samples of 50 individuals each are drawn from the following occu-
pations, and the proportion of women in each sample is ascertained to be as follows:
Sample P, =
1. School teachers 72
2. Nurses 84
3. College professors 40
4. Physician 20
5. Realtors .60
6. Law students a

For each of the six samples, test for significance (nondirectional) the null hypothesis
that the proportion of women in each sample equals the proportion of women in
the population.

Exercise 8.6
A mental health assessment instrument designed to measure a person’s mental
health level on a 30 to 70 scale is known to have a population standard deviation
of 6 = 12. A random sample of n = 25 yields a mean X = 50.
1. Generate a 95% confidence interval for estimating pu.
2. Generate a 99% confidence interval.
Probability Distributions and One-Sample z and t Tests »» 267

3. Suppose o is unknown, but the sample yields a 6 = 11. Generate a 95%


confidence interval.
4. Suppose o is unknown, but the sample’s standard deviation is s = 9. Generate
a 99% confidence interval.
~
Exercise 8.7
A telephone survey of 250 voters shows a local school tax levy passing with 55%
of the vote.
1. Construct a 95% confidence interval.
2. Do the same confidence interval, but assume P, = Q, = .50. What are your
conclusions?
Exercise 8.8
In a training session, a group of managers will be asked to fill out an inventory
designed to evaluate their ability to solve common management problems. The
creators of this inventory have estimated that for the population of all managers,
the mean is 10 and the standard deviation is 3. The possible scores on the inven-
tory range from 0 to 20. Assume normality.

1. What is the probability of randomly selecting an individual with an inventory


score of 15 or above? (outcome A)
2. What is the probability of selecting someone whose score is 7 or below? (out-
come B)
3. What is the probability of randomly selecting someone with a score either 15
and above or 7 and below?
4. What is the probability of selecting someone with a score between 10 and 11?
_ (outcome C)
5. What is the probability of selecting a person whose scores are either 15 and
above, 7 or below, or between 10 and 112
6. If outcome A remains a score of 15 or above and outcome D is the probability
of selecting someone with a score between 10 and 12, what is the probability
of randomly selecting a person whose score is either 15 and above or between
10 and 12?
7. Assume that a score of 15 or above is considered an indicator of a very good
manager (outcome A above). If 43 managers filled out the inventory, what is the
probability of making 2 random selections from this group and obtaining 2 very
good managers?

Exercise 8.9

1. How many different samples of size 3 can be drawn from a population of 6? If


we disregard the order of selection, how many different samples can be drawn?
2. For a population N = 8 and sample size K = 3, calculate the possible number
of permutations and combinations.
268 << STATISTICS FOR THE SOCIAL SCIENCES

NOTES
1. Be aware that there are several other ways of wording the central
limit theorem and the law of large numbers. In addition, these two are
sometimes combined into a single theorem.
2. W. S. Gosset, the expert, published his findings using the pen name
Student. Thus, this test is often called Student’s ¢.
3. Unfortunately, it is hard to find more complete ¢ tables that are relatively
simple to read. Try H. Arkin and R. Colton, eds., Zables for Statisticians (College
Outline Series) (New York: Barnes & Noble, 1963), p. 121, or H. R. Neave, Statistical
Tables for Mathematicians, Engineers, Economists and the Behavioural and
Managerial Sciences (London: Allen & Unwin, 1978), p. 41.
ms
A AEE
wes = ia u bdGs) 9 se) ar . scat) ee

OS ree YE we Pralai Vontives Yoh NAA epegn


er Se Let oe aa a Ve ae Th ws : Uisitec a
gor" : te Al hee hale 2 Ee ’

ZT tye 7 DAs vias :


no wee nh BP nds ae L

¢ A MET Wenham €

Rima ya el a
¥Y KEY CONCEPTS ¥

independent samples/ pooled estimate of Statistical power


independently drawn common variance Type Il error or beta
samples F test for homogeneity error
dependent samples of variances small effects
matched pairs ¢ test/ paired difference ¢ test medium effects
dependent samples research significance large effects
(or paired difference) versus Statistical
t test significance
CHAPTER

Two-Sample ¢ Tests

VY PROLOGUE V¥

With this chapter, we come back from the theoretical and study a family
of tests with widespread research applications. Recall that in Chapter 7’s
prologue, we wanted to study juvenile crime but we couldn’t study every
juvenile criminal. We now know that we can use random samples (which are
small enough for us to study) in place of populations (which are too large
for us to study). So maybe now we have two samples. One is of juvenile
offenders who did time in a detention facility, and the other is a sample of
similar offenders who received probation instead of detention. You as
a researcher have developed an alienation index, which you administer to
everyone in each sample. You then calculate a mean alienation score for
each sample: those in detention and those on probation. Are the sample
differences large enough to conclude differences in the populations?
Another example from Chapter 7’s prologue was the study of married
couples. Suppose you have a group of couples who are having problems
in their relationships and you want to test the efficacy of a particular mar-
riage counseling technique. You take your couples and randomly assign
each couple to one of two groups. One group gets the counseling, and
the other one (the control group) doesn’t. When done, you may compare a
variety of variables to see if there are differences between the two groups,
with (hopefully) the group getting counseling showing improvement in
their interpersonal relationships, as compared to the control group.
MULE ELIEPCRI OY LLLLLECWW<CSHLE LACSEA iia SLO SACI ICY OSCE

pe 271
272 << STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
Like the one-sample ¢ test, the two-sample ¢ test is a comparison of two
means, except that both means are sample means. We no longer know
population parameters, though, as before, we must assume that the popu-
lations are normally distributed along the variable of interest, unless both
samples are large enough to relax these normality assumptions. We com-
pare the two sample means to generalize about a difference between the
two respective population means. The null and alternative hypotheses are
identical to those in the one-sample ¢ test, and H, may be either nondirec-
tional or directional. Since we often do not know or have no basis for esti-
mating the population parameters necessary for the one-sample test, the
two-sample ¢ test is far more commonly used in actual research situations,
where sample statistics alone compose the available data.
The two-sample ¢ test is more complicated than the one-sample variety.
For one thing, this is really a family of tests, and the researcher must select
from this family the formula most appropriate to the data. As we will see,
this selection is based on a variety of factors, such as the way in which the
samples were selected and whether we may assume that the variances of the
populations from which the samples were drawn are equal in magnitude.
Second, in most instances, the two-sample ¢ formulas are more complex
than the one-sample formula and take longer to calculate. Third, there is
one instance in which the degrees-of-freedom formula is so onerous that
many textbooks leave it out altogether or, as in this book, provide the for-
mula but also give an easier-to-calculate approximation of it. Despite these
difficulties, the calculations may still be made in a reasonable period of time,
and because of its widespread usage, it is a particularly important test to
understand.

INDEPENDENT SAMPLES VERSUS DEPENDENT SAMPLES

The first kind of two-sample ¢ test we discuss assumes that there are two
samples (or groups) being compared and that the samples are indepen-
dent; that is, the composition of one sample is in no way matched or
paired to the composition of the other sample. Thus, the two samples
reflect two separate populations. For example, we select a random sample
of 50 men and another random sample of 50 women to investigate gender-
determined views on social issues. Each sample is selected independently
of the other. For each sample, we know its size (7), its mean (X), and its
variance (s°) or, alternatively, its 6°. (For now, we assume we are using s’
rather than 6°.)
Two-Sample t Tests » 273

Independent samples The composition of one sample is in no way matched or


paired to the composition of the other sample. Thus, the two samples reflect two
separate populations.
ose

Sample 1 Sample 2
Size 11, Size 11,
Mean x, Mean x,
Variance s{ Variance s5

Population 1 Population 2
LL, is unknown LL, is unknown
o; is unknown o5 is unknown

Our H) is HL, = U,, and our A, nondirectional is uw, 4 W, (or if H, is directional,


either 1, > LL, or LL, < W,, depending on prior knowledge).
In experimental research, the procedure would be to take a pool of
subjects (or, as they are often called today, participants) and randomly
assign some of the subjects to the experimental group, which will receive
some experimental treatment. The remaining participants will compose the
control group, which will not receive any treatment. The procedure for
selecting the experimental group is identical to drawing a random sample
from the available pool of participants. The random assignment of people to
the experimental group implies that the control group will also have been
randomly assigned. Suppose that out of a pool of 11 people, 6 are chosen to
be in an experimental group that will watch a one-minute television com-
mercial for a well-known product. The format of this commercial has been
used previously with similar products and has a history of raising the view-
ers’ levels of preference for them. Will it work here? The experimental group
watches the ad and then rates the product, assigning a favorability score that
ranges from 0 to 10. The five-member control group rates the product with-
out seeing the advertisement.
The random assignment of the two groups should ensure that only view-
ing versus not viewing the commercial would account for any difference in
favorability scores between the groups. A mean is calculated for each group,
and the two means are compared. If the means differ from one another, is it
due to a real‘difference (as the result of the commercial) or is it due to ran-
dom error or chance? In other words, if the experimental group’s mean is
274 @ STATISTICS FOR THE SOCIAL SCIENCES

higher than the other, there is a possibility that, by chance, more participants
favorable to the advertised product were assigned to the experimental group
than the control group. Thus, the mean of the experimental group could be
higher due to factors other than the commercial they watched.
We expect to find, in any sample that we draw or random assignment that
we make, a certain degree of deviation from the population parameters due
to sampling error. Our test of significance is designed to tell us whether the
differences between the two sample means that we are comparing reflect
a difference between their respective population means—a statistically sig-
nificant difference—or merely reflect the expected sampling variation. In this
case, the population means reflect the hypothetical means that would be
generated if the experiment were to be repeated infinitely. If the latter case
is correct, the observed difference between the two sample means is not
statistically significant, and we do not have enough evidence to conclude
anything other than equality of the two respective population means.
Our experiment is represented as follows:

Group 1 Group 2
Experimental Group Control Group
Size n, = 6 Size n, =5
Mean x, Mean x,
Variance s{ Variance s5

Population 1 Population 2
All People Having Seen the Ad All People Not Having Seen the Ad
LL, is unknown LL, is unknown
o+ is unknown o+ is unknown

Our 1) is 1, = ,, and since we have prior evidence of the success of other


commercials using the same format, our H, is directional: , > ML.
There is another type of research design, used mainly in experimental
research, where the samples are dependent. Members of one sample are
not selected independently but are instead determined by the makeup of
the other sample. We call this a matched pairs situation, and another type
of ¢ test, a dependent samples f¢ test, is used. To understand what we
mean by a matched pair, imagine a situation where we start with a set of
pairs of identical twins. One of each pair of twins is assigned to the experi-
mental group, and the other member of that pair goes to the control group.
Two-Sample t Tests p 275

Thus, being included in the control group is dependent on one’s twin being
included in the experimental group. Then, presumably, each group would
share identical inherited traits, making any difference between the groups a
function of environmental (as opposed to hereditary) differences—notably,
the effects of the experiment. Sometimes, the pairs are not twins but are
related in other ways (for example, wives and their respective husbands). Or
the pairs could be based on other factors such as age and race (for example,
if one group includes a 35-year-old Caucasian female, the second group
would include another 35-year-old Caucasian female),

Dependent samples or matched pairs Situation in which members of one sample


are not selected independently but are instead determined by the makeup of the
other sample.

Dependent samples t test The t¢ test used when the two samples are dependent
samples.

A very common matched pairs situation is a before-after or repeated-


measures experiment. Imagine that each member of a panel was asked to
rate each of two candidates contending for the same elective office. Later,
the panel watches a 2-hour debate between the contenders. At the end of
the debate, each panel member rates the candidates again. The matched
pair is each person’s “before” score with that person’s “after” score. Any dif-
ference would presumably be due to the debate watched. We will return to
this subject later when the dependent samples ¢ test is discussed.

THE TWO-SAMPLE ¢ TEST


FOR INDEPENDENTLY DRAWN SAMPLES

In this instance, there is no matching but rather two independently drawn


samples or randomly selected groups. The sampling distribution of which
we make use is the ¢ distribution, except it is derived from the sampling dis-
tribution of the differences between all theoretically possible pairs of sam-
ple means. (Consult a more advanced text for more details.) As with the
one-sample ¢ test, the difference between our two sample means is divided
by the standard deviation of the sampling distribution, the standard error, to
find t. Now, however, our standard error is the standard deviation of sample
mean differences, a different standard error than the 6/Jn or the S//n-1
used in the one-sample ¢ test.
276 « STATISTICS FOR THE SOCIAL SCIENCES

This raises a new complication: The standard error we seek is an estimate


based on the variances of our two samples. If the two sample variances are
close in magnitude, we may assume that both parent populations from which
the samples were drawn have the same variance, that is, 0;= 05. If these pop-
ulation variances are the same, we can calculate the standard error of our ¢
formula using what we call a pooled estimate of common variance. This
is based on a weighted average of our two sample variances being used to esti-
mate the population variance in finding the standard error. If we can make use
of this pooled estimate, we generally will have a greater opportunity to reject
the null hypothesis than would be the case when equal population variances
cannot be assumed. If, however, we cannot assume equal population vari-
ances, that is, 0; # 05, we must use a different formula for finding the standard
error and, thus, a different ¢ formula. To determine which ¢ formula is most
appropriate, we can use the F test for homogeneity of variances. (The test
is shorter than its title!) Computer programs that run ¢ tests will usually run
both the equal and unequal population variance ¢ formulas and also do an F
test to help the reader ascertain which obtained ¢ value is more accurate.

Pooled estimate of common variance Estimate based on a weighted average


of two sample variances being used to estimate the population variance in
finding the standard error.

F test for homogeneity of variances _A test, based on the sample variances, used
to determine the most appropriate t test formula to use.
SAS
HER LEN LOO

To avoid confusion, let us stop here and lay out a set of steps for the
two-sample ¢ test (independent samples) and then illustrate these steps with
an example. At the appropriate points, the formulas will be given and
explained for the F test for homogeneity of variances, the equal population
variance ¢ test, and the unequal population variance ¢ test. First, the steps:

1. Write out H, and 7, for the original problem, the comparison of the
two sample means.

2. For each sample, determine its 7, its ¥, and its variance, s?.

3. To determine which ¢ formula to use (that for equal population vari-


ances or that for unequal population variances), do the F test for
homogeneity of variances.
a. Write out H, and H, for the F test (these are not the same as the
ones for the ¢ test).
b. Calculate F and its two degrees of freedom.
Two-Sample t Tests p 277

c. Compare the obtained F to F critical? .05 level, taken from a table of


critical values of F
Cea obtained
2
— critical? assume unequal population variances.
If Fobtained <F critical? assume equal population variances.

4. Perform the appropriate ¢ test as determined by the F test.

Example 1. Let us work through these steps, using the example of the TV com-
mercial. Recall that 6 people (the experimental group) will see the commercial
and then evaluate the product’s favorability. The other 5 (the control group)
will evaluate favorability without seeing the commercial. Remember also that
because of previous success with the same format for the ad, we expect favor-
ability to rise once the viewing is complete, so our H, is directional. Note also
that because of the smallness of our groups (72, = 6 and 7, = 5), we must
assume that favorability is normally distributed in our two populations.

1. Write out H, and H;:

(Group 1 is the experimental group.)

2. Determine 7, X, and s’ for each sample. Following are the data:

Sample 1 Sample 2
Saw Commercial Did Not See Commercial

8
il x,=
8
3
5
6
7
= ON US
uleesrentres
Wears

M4ma) \| IN~S >


5x,
= 29

= esi 47 Z y
xXy= ee
1} 6 n12 5

To get the sample variances using the computational formula, we must


square our two sets of scores. (Later on, we will ease your burden by pro-
viding you with the variances in advance.)
278 a STATISTICS FOR THE SOCIAL SCIENCES

Xi xX, =

100 64
36 9
64 25
49 36
81 49
Bid Senet
Sie = p12, SO Slee:

x1)? » ON x2)"
oa xi ai a oe Ye n2
oe ny i Ss

Ciae are6a age MiP: eter


(47)2

5
wy

379 — 2206 183 — &


Moe a 5
_ 379 — 368.17 _ 183 — 168.2
a 6 is 5
= thot! = 296

Summarizing our results:

Sample 1 Sample 2
Saw Commercial Did Not See Commercial
n,=6 nN ,=5
Xx,=7.83 x, = 5.80
= sl S200

3. We perform the F test for homogeneity of variances.


a. In the F test we are doing, the null hypothesis is that the popula-
tion variances are equal, and the alternative hypothesis is that they
are not.

ts = Oe
ae (ane x ois

(In the F test, H, is always nondirectional.)

b. To calculate F divide the larger of the two sample variances by the


smaller one.
2
larger
r= 22
52
smaller
Two-Sample t Tests j» 279

In this case, the larger s* is the one for Group 2, the control group. Thus,

Fobtained = > = eee


Ss 1.81 ;

The F has two degrees of freedom, one associated with the numerator and
one associated with the dendminator. In each case, we subtract one from the
sample size. The numerator degrees of freedom is one less than the size of the
sample having the larger variance. The denominator degrees of freedom is
one less than the size of the sample having the smaller variance. In this case,

df numerator =5-1=4
df denominator = 6-1=5

c. Our obtained F is then compared to F.,,,,.4, 05 level, df = 4 and 5.


Table 9.1, a portion of a fuller F table to be presented later, gives
F
siticats FO the .05 level only. In the table, 7, means the numerator
degrees of freedom. (7, and, here mean degrees of freedom, not
sample sizes!) We move along the row until we find the column for
the appropriate degrees of freedom. In this case, 7, = 4. We go
down the 7, column until we find the row for our degrees of free-
dom, 7, = 5. We move along the 7, = 5 row until it intersects the
n, =4 column. The number at that intersection, 5.19, is OU Fjirica’
[Link] rejecithe A, torthe F test, Fs cineq MuUstrequal or‘exceéd this
Eien DOUU spay ae OF ANOLE ga 1D Fohiney = MOL <I ae,
.05, (df= 4 and 5) = 5.19. We cannot reject H,, so we now perform
the ¢ test designed for equal population variances, to be presented
in Step 4.

It should be noted here that many computer routines, including two of


those to be discussed shortly, calculateF from 6? rather than s*. Had we done
that here, our obtained F would have been 3.700/2.166 = 1.708, still less than
Focicar Lhe difference between 1.708 and the 1.64 above would have been
smaller if our sample sizes had been larger, which is almost always the case
in nonexperimental research. Rarely would the discrepancy between the two
ways of obtaining F change our decision as to which ¢ test to use.

4, Since equal population variances may be assumed, we use the follow-


ing formula in which the denominator is the pooled variance estimate.

ipa
==
nist +1285 1 ae Alls
ny—n2—2 ny n2

df =n, +n2-2
280 <€ STATISTICS FOR THE SOCIAL SCIENCES

Table 9.1 Critical Values ofF (.05 level only) for the F Test for Homogeneity
of Variances

n\n, i Z te; 4 5 6 8 12 24

1 161.40 199.50 215.70 224.60 230.20 234.00 238.90 243.90 249.00


2 18.51 19.00 19.16 19.25 19.30 1935 19.57 19.41 19.45
2) 10.13 Eas) 9.28 Dalle 9.01 8.94 8.84 8.74 8.64
4 tor 6.94 6.59 6.39 6.26 6.16 6.04 5.71 =
ea
5 6.61 19 5.41 Bly, 5:05 4.95 4.82 4.68 4.53
6 Sy) 5.14 4.76 eo) 4.39 4.28 4.15 4.00 3.84
7 Srey. 4.74 Cw) 4.12 D7 3.87 STD DoT 3.41
8 5:32 4.46 4.07 3.84 3.69 3.58 3.44 3.28 S12
9 5.42 4.26 3.86 3:63 3.48 DOT P29 3.07 2.90
10 4.96 4.10 371 3.48 3.33 3.22 3.07 S01" 274
is 4.84 3.98 359 3.36 3.20 3.09 2.95 2.79 261
12 4.75 3.88 3.49 3.26 3.11 3.00 2.85 209. 250
13 4.67 3.80 341 3.18 3.02 2.92 277 560 242
14 4.60 3.74 oe 2.96 2.85 2.70 253. 2.35
15 4.54 3.68 3.29 3.06 2.90 2.79 2.64 248 2.29
16 4.49 3.63 324. © 3/04 2.85 2.74 2.59 Map. 234
17 4.45 3.59 20h 2.96 2.81 2.70 2.55 238 2.19
18 4.41 3.55 Buln unos 27 2.66 2.51 284-) “255
19 4.38 3.52 313 2.90 2.74 2.63 2.48 gor 31
20 4.35 3.49 aoe |) 287 2.71 2.60 2.45 2.28 2.08
21 4.32 3.47 3.07 —-2.84 2.68 O57 2.42 2.25 2.05
22 4.30 3.44 3.05 2.82 2.66 2.55 2.40 gos: 3.08
23 4.28 3.42 3.03 —-2.80 2.64 2.53 2.38 220 2.00
24 4.26 3.40 3.01 2.78 2.62 2.51 2.36 2.18 1.98
25 4.24 3.38 20 3G 2.60 2.49 2.34 2.16 1.96
26 4.22 3.37 298 2.74 2.59 2.47 2.32 2.15 1.95
27 4.21 3.35 296 2.73 2.57 2.46 2.30 2.13 1.93
28 4.20 3.34 205 = 5 2.56 2.44 2.29 2.12 1.91
29 4.18 3.33 2.93 2.70 2.54 2.43 2.28 pig 456
30 4.17 3.32 202 "365 2.53 2.42 ey, 2.09 «1.89
40 4.08 3.23 2.84 2.61 2.45 2.34 2.18 2.00 1.79
60 4.00 3.15 276. 250 237 2.25 2.10 1.92 1.70
120 3.92 3.07 2.68 2.45 2.29 aly, 2.02 183 1.61
ms 3.84 2.99 260 ey 2.21 2.09 1.94 1.75 1.52
SOURCE: Abridged from ‘Table V of R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricultural
and Medical Research (oth ed.), 1974. Reading, MA: Addison-Wesley, an imprint of Pearson
Education.
NOTE: Values of 7, and 72, represent the degrees of freedom associated with the larger and smaller
estimates
of variance, respectively. p = .05.
Two-Sample t Tests ® 281

It’s easiest to first do the components of the ¢ formula and then put
them together.

Nix, = 1,0) — 5.50 = 2.05

mist +m2s> _ 6(1.81) +5(2.96) 10.86+14.80 25.66


ny +nz-2 ee ee ee
BiPree
eeenye eae:O17 +020
+0, 20 = 0:37

Nis? eites)
(p ; (1 1] —
+ = (2.85) (0.37)
ny +n2—2 nN\

= 7 1.0545 = 1.0268 = 1.03

Thus,

Fhe Oe 2.03
(— 19708 = 197 1
118+ +285 1 1 1.03
nm, +n2—2 ny ats n2

lobtained = LOT

The df=n,+n,-2=6+5-2=11-2=9. We use the critical values of


t for a directional H,. See Table 9.2.

At .05 level, 7_.4- @/ =9) = 1.833 <.1.971 reject H,


At .025 level, tonic (Gf=9) = 2.262 > 1.971 p<.05
Note that if H, had been nondirectional,

At .05 level, t..)critical f= 9) = 2.262 > 1.971


and we would not have been able to reject H,.
Since H, has been rejected, we conclude our alternative hypothesis of
H, > uw. As in previous instances, this particular commercial resulted in
increased favorability ratings for the product featured in the commercial.
Since H, has been rejected, we continue to assume that if the entire con-
sumer population had viewed the advertisement, their mean support score
LL, would also increase.
Remember that there are risks in using directional alternative hypothe-
ses. Not only do we need prior evidence of the assumed direction, as
282 < STATISTICS FOR THE SOCIAL SCIENCES

Table 9.2 Distribution of ¢

Level of Significance for One-Tailed Test

10 OS 025 OL 005 0005


Level of Significance for Two-Tailed Test

df 0 10 05 02 ror 001
1 3.078 6.314 12.706 31.821 63.657 636.619
2 1.886 2.920 4,303 6.965 9.925 31.598
3 1.638 2.353 3,182 4.541 5.841 12.941
4 1.533 2.132 2.776 3.747 4.604 8.610
5 1.476 2.015 2573 3.365 4.032 6.859
6 1.440 1.943 2.447 3.143 3.707 5.959
i 1.415 1.895 2.365 2.998 3,499 5.405
8 1.397 1.860 2.306 2.896 3.355 5.041
oe 1.383 1.833 2.262 2.821 3.250 4.781
10 1.372 Sie 2.228 2.764 3.169 4.587
11 1.363 1.796 2201 2718 3.106 4.437
12 1.356 1.782 2.179 2.681 3.055 4.318
13 1.350 ez 2.160 2.650 3.012 4.221
14 1.345 1.761 2.145 2.624 2.977 4.140
15 1.341 1.753 2.131 2.602 2.947 4.073
16 1437 1.746 2.120 2.583 2.921 4.015
17 1.333 1.740 2.110 2.567 2.898 3.965
18 1.330 1.734 2.101 2552 2.878 3.922
19 1.328 1.729 2.093 2.539 2.861 3.883
20 1.325 1.725 2.086 2.528 2.845 3.850
21 1.323 1721 2.080 2.518 2.831 3.819
OD 1.321 leralyy 2.074 2.508 2.819 3.792
23 1.319 1.714 2.069 2.500 2.807 3.767
24 1.318 lepine 2.064 2.492 2.797 3,745
25 1.316 1.708 2.060 2.485 2 787 3.725
26 1.315 1.706 2.056 2.479 2.779 3.707
27 1.314 1.703 2.052 2.473 2.771 3.690
28 1,313 1.701 2.048 2.467 2.763 3.674
29 1.311 1.699 2.045 2.462 2.756 3.659
30 1.310 1.697 2.042 2.457 2.750 3.646
40 1.303 1.684 2.021 2.423 2.704 3551
60 1.296 1.671 2.000 2.390 2.660 3.460
120 1.289 1.658 1.980 2.358 2.617 3.373
00 1.282 1.645 1.960 2.326 2.576 3,291
SOURCE: Abridged from Table V of R. A. Fisher and F. Yates, Statistical Tables for Biological,
Agricultural and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley,
an imprint
of Pearson Education.
Two-Sample t Tests p 283

was the case here, but we must always make sure that our findings are
consistent with the direction assumed. Suppose Xx, had been 3.77 instead
of 7.83, but we retained the MU, > kt, alternative hypothesis. The numerator
of the ¢ formula would be 3.77 — 5.80 = —2.03, and ¢ would be —1.971. Using
its absolute value of 1.971, we would reject H, in the same way as just done.
Yet, ifanything, our two sample means suggest an HT, of UW,< M,, and we have
findings inconsistent with the original H,. Even though we reject H,, we
cannot conclude U, > U,. In this case, our only option would have been a
two-tailed H,, but we have already seen that, in such a case, we cannot
FEJECE F7.
Note, finally, that it is conceivable—albeit improbable—that really
Ll, > LH, but, due to sampling error, X, < X,. However, we have no basis for
knowing that fact when we do our study. Accordingly, if ¥, <x, but H, said
Jt, yf, OF the reverse, x. > x, butud,, Stated Uu.-< u-do 208 proceed with a
one-tailed test.
To summarize to this point, we established H, and H, for our data; found
n, X, and s* for each sample; and did the F test for homogeneity of variances
to determine the appropriate ¢ formula to use. In this case, the F test led us
to use the ¢ formula where equal population variances are assumed. Using
the appropriate ¢ test, we were able to reject H, with a probability of p < .05
and conclude that in the population, viewing the commercial enhances
support for the product featured. Now, let us see what would happen if the
F test concluded unequal population variances.

Example 2. Suppose Sample 1 remained the same, but the scores for
Sample 2 were as follows:

X=
10
i!
)
6
n,=5 =p)
SH 29)

Be AT
12 5

The sample size and sample mean stay the same as before, but the
sample variance is now larger.
284 @ STATISTICS FOR THE SOCIAL SCIENCES

36
81
yx?
=227
Thus,
yx? — as 227 — ee 227 — 84)
2 —

oa n2 a 5 5

a 227 — 168.2 = 58.8 ~ 11.76


> 5

Redoing the F test for homogeneity of variances,

_ stul
Sg 176
181
_6oy ae
d

AO = 0;
At df= 4 and 5,
Poy OS=5.10 G97" iwejectii “p< 05

We may not assume equal population variances.


Once again, if Fojrainea Were Zenerated using the 6, we would find that
F = 14.700/2.166 = 6.786. Just as above, we may not assume equal popula-
tion variances.
Where population variances are assumed to be unequal, we use > the
following formula:

As before, we first calculate the components.

Xp = Ny = /-05, — 5.60 = 205

5 2Bee ee
181 181
oe 0.36
Two-Sample t Tests 285

& TIGA
es = 24
Vee Seat ee

Using as df the smaller value of 7, or 7,, df= 5.


Ror ardirectionalii7, #2.) Od evel, @f=5)= 2.015 >.1. 115. Wecannot
rejectl.,
The degrees of freedom for this problem, the lesser of 7, or 7,, is
only a substitute for a complex degrees-of-freedom formula used by
packaged statistical computer programs. This larger formula is a more
accurate approximation but is often too cumbersome for noncomputer
applications.

If we did use it in the above ¢ test, we would also work it in stages.


2
2 2 z

pete ole 5 ny—1

2 6 st ‘
ante ee LG, 2_ |— (2.94)?
= 8.64
i eae nz—1
LZ
§ a 3 | 2.94)? ==(3.30)*
= (0.36 + 2.94)* (3.30)? = 10.89
ny —1)

(4s) ee
s2 :

Gia iyes=t% 5
3 \?
(4) tS 8.64 A 8.64 =F 16

(w2-1) 5-1 4
286 << STATISTICS FOR THE SOCIAL SCIENCES

Thus,

(ah ie a). 10.89 _ 10.89 _ ¢ o


a (2 ) i sae ee
nz—1

ny—1 n2—1

A computer would list df as 4.97.


If we were doing this by hand, we would round down (not up) and use
df = 4 in our ¢ table. Again, this makes it harder to reject H, than it would
have been had we been able to round up to 5 degrees of freedom. We get
essentially the same results as before: ¢,,,,,.,, one-tailed, .05 level, (df = 4) =
2.132 > 1.115, and we cannot reject H,. The ad has no effect on favorability
scores (in this example, as modified for the assumption of unequal popula-
tion variances).

ADJUSTMENTS FOR SIGMA-HAT SQUARED (<”)

As was the case with the one-sample ¢ test, it is possible that instead of s?
for each sample, 7 — 1 replaced m in the denominator of the formula, and
consequently, what was calculated was 6’, not s*. We then would need to
modify our ¢ formulas accordingly.
In the first example, where population variances were assumed

Sample 1 Sample 2
n= n,=5
X= 785 x, = 5.80
OF = 2.17 65 = 3.70

We would make the following modification in the ¢ formula for equal


population variances:

. df=n,+n,-2
PC Wee69) if J ie
ip =11-2
(11 -1)G2-+(n2—1)67 ]F 4 1 =9
/| nN -+n2—2 E ty nz

As before, X,—*X, = 2.03 and 1/n, + 1/n, = 0.37. Recalculating the remaining
expression to adjust for 6%,
Two-Sample t Tests p 287

(m1 — 1)6f + (2-167 — 6©— 192.17) + (5 — 1)G.70)


Ni 1 — 2 - 6+5-2

_ 5(2.17) + 4G.70) _ 10.85 + 14.80


i=2 ‘¢ 9
25,05
= 9 — = 2,0 5

just as it was in the original formula using s?.

= 2.03 _ 2.03 _ 2.03 _, a


J/(2.85)(0.37) 1.0545 1.03
a—s Oral

In the example where unequal population variances were assumed, we


would find:

Sample 1 Sample 2
n,=0 n,= 5
X, = 7.83 X, = 5.80
Coy 6? =14,7
We modify the ¢ formula for unequal population variances as follows:

: Mie af =the lesser of 7, or 72,

X,—X, is still 2.03. Recalculating the denominator to adjust for oe,

ee DAG NAG.
“1 a - Shey gee a
1 2

So, t = 2.03/1.82 =1.115 as it was when we used sj and s>.


Finally, the alternative longer degrees-of-freedom formula for unequal
variance ¢ tests would become
As before, df = 4.97
288 << STATISTICS FOR THE SOCIAL SCIENCES

INTERPRETING A COMPUTER-GENERATED ¢ TEST

Figure 9.1 presents the SPSS T-TEST printout for Example 1 in this chapter.
Starting on the left of Figure 9.1, some general statistical information is
presented. The dependent variable has been coded VAROOO02 by the
researcher. For each group (category of VARO0001), the printout lists its
size, mean, standard deviation (SPSS uses the 6? formula and not s*), and
standard error. Note the box below labeled “Independent Samples Test.”

Figure 9.1 = SPSS Printout for Example 1

t-lTest
SP

Group Statistics

Std. Std. Error


VAROOOO] Deviation Mean

VAROO0002 1.47196 .60093


1.00

1.92354 86023 |

Independent Samples Test

Prey test for Equality of Means

Test for 95% Confidence


Equality of Interval of the
Variances t-test for Equality of Means Difference

f Lower | Upper

VAROOO02| . .O7E 4.34508


Equal
variances
assumed

Equal
variances
not
assumed
Two-Sample t Tests p 289

Instead of using the F test for the homogeneity of variances presented


in this chapter, SPSS uses Levene’s test for equality of variances, which is less
dependent on the normality assumption than the F test for homogeneity of
variances presented earlier in this chapter. We interpret the Levene findings
exactly the same way as the F test. Since Levene’s F of .258 generated a prob-
ability of .623 (Sig.), greater than .05, we cannot reject a null hypothesis of
equal population variances.
Accordingly, we will use the ¢ test, with equal variances assumed. Under
the title “¢ Test for Equality of Means,” note the ¢ value of 1.990. That is the
one we want. (Below it is 1.938, which is ¢ when equal variances are not
assumed.) To the right of the ¢ values are the degrees of freedom. To the
right of degrees of freedom, under Sig. (two-tailed) is the probability—the
exactp value (not a < .05 statement but the actual probability). The 4 df and
p values on this line are all based on the formula for equal population vari-
ances. On the line below, we find the 4, df andp values from the formula that
assumes unequal population variances. The degrees of freedom for the
unequal population variances ¢ test is based on the long formula. Recall that
for Example 1, we calculated the ¢ only for equal population variances. Our
calculated ¢ of 1.971 differs from SPSS’s 1.990 since we used fewer decimal
places in calculating ¢, Our df and SPSS’s are the same, 9. The probability on
the printout is the two-tailed probability .078, not the one-tailed probability
that we used in our hand calculations. To get the one-tailed probability,
divide .078 by 2. Thus, the exact directional probability is .039—less than .05
and greater than O01.

1. Examine the probability of F.


a. Ifthe probability of Fexceeds .05, we cannot reject a null hypoth-
esis (for the F test) of equal population variances. Use the f¢ test
results to the right of “equal variances assumed.”
b. If the probability of F is less than or equal to .05, we reject the
F test’s H, and assume unequal population variances. Use the
t test results to the right of “equal variances not assumed.”

2. Examine the appropriate (EQUAL or UNEQUAL) probability of ¢.


a. If the probability (Sig.) of |7| exceeds .05, we cannot reject the
initial null hypothesis (the one for the ¢ test). The difference is not
statistically significant, Md, = L,.
b. If the probability of |7| is less than or equal to .05, we reject the
¢ test’s null hypothesis and conclude H,,. (Note: If H, is directional,
divide the significance by 2 before working Step 2.) For this spe-
cific problem, the probability of F 1623, exceeds .05. We cannot
290 << STATISTICS FOR THE SOCIAL SCIENCES

reject H,: 0? = 03, so we will use the ¢ test for equal population
variances. Looking along the “Equal Variances Assumed” row, we
find a probability (significance) of .078. Since our original H, was
one-tailed, divide the .078 probability by 2. The result is .039 (as we
previously demonstrated). Since .039 is less than .05, we reject our
original H, of U, = Mand conclude H;: M, > Lb.

COMPUTER APPLICATIONS:
INDEPENDENT SAMPLES ¢ TESTS

Let us take the data from Example 1 upon which Figure 9.1 was based, set
it up, and run the two-sample ¢ test using SPSS. We will then do the same for
SAS and Excel. Before starting, you may want to review the setup instruc-
tions for SPSS and SAS presented in Chapter 6. (Excel was not presented
then because it has no current routine for crosstabs.) We will also compare
the outputs from the three programs.

SPSS

Variable 00001 will be whether or not the respondents saw the com-
mercial, coding 1 if they saw it and coding 2 if they did not see it. Variable
00002 will be the favorability rating. Table 9.3 shows the data list. We then
click on the following menu options:

Table 9.3

VAROOOOL VAROO002

] 1.00 10.00
2 1.00 6.00
3 1.00 8.00
4 1.00 7.00
5 1.00 9.00
6 1.00 7.00
of 2.00 8.00
8 2.00 3.00
y) 2.00 5.00
10 2.00 6.00
11 2.00 7.00
Two-Sample
t Tests » 291

Analyze
Compare Means
Independent Samples
In the Dialog box, highlight VARO0002 and use the top button with the
pointer to click it over to the Test Variable box. We then highlight VAROOOO1 and
move it into the Grouping Variable box, using the lower button with the pointer.
We then click the define groups button and type 1 to the right of “Group 1” and
2 to the right of “Group 2.” Then click continue, bringing us back to the first
Dialog box. Then click ok to get the output presented in Figure 9.1.

SAS

At the top of the screen, click as follows:

Solutions
Analysis
Analyst

While it is possible to list the data the way it was done for SPSS and run
the same SAS ¢ test routine, it is easier to code both SAS and Excel differ-
ently. In column A, enter the six scores for Sample 1 (saw the commercial),
and in column B, enter the five scores of the Sample 2 members (didn’t see
the commercial) (see Figure 9.2).

Figure 9.2. SAS Data and Printouts for Example 1

Input:

A B Cc ID)

1 10 8
2 6 =)
3 8 5
4 y 6
5 9 Z
6 y)
=

Output:
Two-Sample Test for Variances of B and A

(Continued)
292 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 9.2 (Continued)

Sample Statistics

Hypothesis Test

Null
hypothesis: Variance I / Variance 2 =1
Alternative: Variance 1 / Variance 2 ~=1

— Degrees of Freedom —

Two-Sample ¢ Test for the Means of A and B


Sample Statistics

22295 0.0009

E9255
=| 0.8602

Hypothesis Test

Null
hypothesis: Mean 1 —Mean 2 =0
Alternative: Mean 1 —Mean 2 ~=0

If Variances
Are t statistic te ae a

Equal 1.990 0.0778

Not Equal 1.938 0.0914


Two-Sample t Tests p 293

At the top of the screen, click

Statistics
Hypothesis Tests
Two-Sample Test for Variances

This is because with SAS, the F test for homogeneity of variances is run
separately from the ¢ tests. This is similar to the F value we calculated earlier
in this chapter. In the upper left portion of the Dialog box, you will see
Groups are in and, below it, one variable with two variables below that.
Click on two variables. Below, just above the word remove, you will see a
box with A and B in it. You will need to highlight (left click) each letter and
move it to either group 1 or group 2 by clicking on the appropriate button.
Here you must decide whetherA should be moved to group 1 or group 2.
Remember, we must divide the /arger variance by the smaller! Now we
know from our earlier calculation that Sample 2 (B) had a variance of 3.700,
and Sample 1 (A) had a variance of 2.166, so B, with the larger variance,
should go into group 1, andA, with the smaller variance, should go to group
2. What if we didn’t know the variances? We could use the descriptive sta-
tistics SAS routine, but there is an easier way. Click B into group 2 and A into
group 1. Then click the ok button.
The first output reproduced in Figure 9.2 appears. Note under Sample
Statistics that B, with the larger variance (3.7), appears first and A, with the
smaller variance, appears under B. A’s variance, 2.166667, appears directly
under B’s variance. If the larger variance is not above the smaller one, redo
the test, reversing the letters in group 1 and group 2! Or simply examine F
If it is less than one, reverse the groups and redo the run, or recalculate F
by hand from the variances on the printout.
Below, where it says Hypothesis Test, you will note that F = 1.71, degrees
of freedom are 4 and 5, and the exact probability is 0.5074, way above .05,
sO we cannot reject the null hypothesis of equal population variances. Now
we run the f test:

Statistics
Hypothesis Tests
Two-Sample ¢ Test for Means

The Dialog box resembles the one you just had for the F test. Click the
groups are in to the two variables category. Move A to group 7 and B to
group 2. (If you reversed the groups forA and B, all it would do is change
the sign of ¢ from positive to negative.) Click ok.
294 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 9.2 also displays the ¢ test run. (Don’t be concerned about the
way the null and alternative hypotheses are stated; they mean the same as
what we have been using.)
You see that the ¢ statistic for equal population variances is 1.990, df= 9,
and the probability (two-tailed) of 0.0778 must be divided by two for our
directional alternative hypothesis, yielding 0.0389, less than .05. (If you
don’t feel like dividing, you could have moved the alternative hypothesis in
the Dialog box to mean 1 — mean 2 < 0.) So we reject the null hypothesis
for the ¢ test. The commercial had the desired effect.

Excel

Make sure your Excel program has the Data Analysis Toolpack Add-ins
installed. When you open the program, a spreadsheet appears. Note at the
bottom left-hand side that it says sheet 7. The rows are numbered, and the
columns are labeled with letters of the alphabet, just as with SAS. The data
are entered exactly as they were with SAS, resembling the way the data for
Example 1 first appeared in this chapter, with each column on the spread-
sheet being a different group. Column A contains the scores for Sample 1,
those who saw the commercial, and column B contains the scores for
Sample 2, those who did not see the commercial. The data entry pattern and
the output subsequently generated will all be found in Figure 9.3.
To first do the F test for homogeneity of variances, note the button at
the top labeled tools and click as follows:

Tools
Data Analysis
F Test Two-Sample for Variances

A Dialog box appears in the center of the screen. The cursor will be
in the left side of a box labeled Variable 1 Range. You can actually left
click your mouse on column A, number 1 and, holding the button down,
move to column A, number 6. This highlights all the data in column A..When
you lift the left mouse button, in the Variable 2 Range box, you will see
$A$1:$A$6, telling you that the variable range is from column A, row 1 to
column A, row 6, (You could also have simply typed that information into
the box, instead of clicking it in with your mouse.) Now click in the box
above the first one, which is labeled Variable 1 Range, and go to column B
with your mouse and left click, highlighting column B, number 1, down to
column B, number 5. When done, you will see $B$1:$B$5 in the Variable 1
Range box. Remember, just as before, we want the variable in the Variable 1
Range box to be the variable with the /arger variance. Click ok, and the out-
put will appear in the upper left-hand portion of the screen (see Figure 9.3).
Two-Sample
t Tests p 295

Figure 9.3 Excel Data and Printouts for Example 1

il 10
2 6
3 8
4 #
5 9 1:00
499)
AON
GN
SI

6 ri
5

Sheet 1

F-Test Two-Sample for Variance

Mean axe) Us oee see)

Variance Sef 2.166667

Observatio 5 6

df 4 5

F 1.707692

P(F <=f) on 0.283722 |

F Critical o [ 5.192168 |

Sheet 4

(Continued)
296 4 STATISTICS FOR THE SOCIAL SCIENCES

Figure 9.3 (Continued)

t Test: Two-Sample Assuming Equal Variances

Variable 1 Variable 2

TO99999 5.8

Variance 2.166667 eh,

Observatio 5

Pooled Var

Hypothesiz

df 9

P(T<=t) tw 0.077832

Sheet 5

t Test: Two-Sample Assuming Unequal Variances

Variable 2

Variance

Observatio

Hypothesiz

df

L.337729

0.040924

t Critical or 1.894579

P(T<=t) tw 0.093848

t Critical tw 2.364624

Sheet 6
Two-Sample t Tests j» 297

Again, confirm that the variance on the printout for Variable 1 is larger
than the variance for Variable 2. If it is not, redo the run, reversing the range
boxes of the two variables. Since F is 1.707692, and its probability is greater
than .05, we cannot reject the null hypothesis for equal population vari-
ances, SO we want to do the ¢ test that assumes equality of variances.
Looking below and to the left, note that your F output is on Sheet 4. Click
on Sheet | to return to the page where you keyed in the data.
Follow the same procedure as you did for the F test, but call up the
appropriate ¢ test instead:

Tools
Data Analysis
t Test: Two-Sample Assuming Equal Variances

The Dialog box is identical to the one for the F test. This time, drag
column A and move the information into the Variable 1 Range box. Click the
cursor to the Variable 2 Range box, highlight column B, and move that infor-
mation into that box. If, as before, you put column B in the Variable 1 Range
box and A in the other, all you would do is change the sign of¢ from positive
to negative. (No thinking outside the box!) Click the oR button as before.
You are now on Sheet 5, and your test results, as before, are in the upper
left portion of your screen. Note that the ¢ is 1.989718 and the probability is
0.077832. Again, remember that we divide the probability by 2 because our
t test’s initial alternative hypothesis was directional.
The other ¢ test, for unequal variances, is not needed in this application,
but just to complete this exercise, let’s do it anyway.

Tools
Data Analysis
t Test: Two-Sample Assuming Unequal Variances

Handle the Dialog box exactly as before. Click ok. The output on Sheet
6 is reproduced along with the others in Figure 9.3.

THE TWO-SAMPLE t TEST FOR DEPENDENT SAMPLES

In the case of a before-after experiment or a situation where members of


one group are matched to specific members of another group, the samples
are not independent, and a separate ¢ test must be performed. While we
call this test a two-sample ¢ test for dependent samples (or matched pairs),
actually the test is a one-sample f test applied to the differences in each pair
of scores. Accordingly, the test is also called a paired difference ¢ test.
298 << STATISTICS FOR THE SOCIAL SCIENCES

Paired difference t test. A one-sample t test applied to the differences in each


pair of scores.

Suppose in Example 1 that before viewing the commercial, our


6 subjects rated the product (before scores), then saw the ad and rated the
product again. The before and after scores are presented below.

Before After
5 10
5 6
7 8
8 7,
7 9
3 7

For our sample of 6 people, there are 6 pairs of scores, 7, = 6, where 72,
is the number of pairs. The first subject listed scored 5 before and 10 after
seeing the ad.
Note that in all but one case, favorability rose after viewing the com-
mercial. The mean favorability score rose from 5.83 (before) to 7.83 (after
the viewing). We are trying to demonstrate that viewing the commercial
causes the mean favorability to change. Based on previous positive results
with this same format, we expect the change to be an increase in favorabil-
ity. Thus, we use a one-tailed H,.

Fy: Hpetore a Hater

HT: Peeton s Psteer

To what extent are we safe in assuming that the increase for our 6
subjects reflects an increase among the population of all people who would
view the commercial? Since the test we will be performing deals with the dif-
ferences in observed scores, the null and alternative hypotheses could be
written in terms of these differences.

Ay: Mp = 0
ig Ba EN

By L, we mean the mean of all the differences in scores in the population.


If we subtracted everybody’s before score from everybody’s after score, we
Two-Sample t Tests PB 299

would have a difference score for everyone. If a score goes up from before
to after, the difference will be a positive number; if the score goes down, the
difference will be negative. Since we expect the scores to rise, our direc-
tional H, is 4, > 0. If we had made no directionality assumption, our H,
would be u,, # 0.
We use each difference of scores in our sample (D) as the basis for this
Etest.
D- HD
t
a Sp//Mp — 1

where

n, is the number of matched pairs of scores in the sample.


D is the difference (after score — before score) for each pair of scores in
the sample.
D is the mean of all our sample’s difference scores (retaining their +
or — signs).
S, is the sample standard deviation of the difference scores:

EOS
Sp=
ge

ll, is the mean of the difference scores for all possible pairs in the
population. As per the null hypothesis, UW, is assumed to equal to zero.
Accordingly, we may delete it from the formula.
Thus, =
18)
SS
SD/a/fip aA

where

> (D-D/
p=
Mp

We first find S,,, following these steps:

1. Subtract each before score from each after score to get each D.
2. Add the Ds algebraically to get )~D and divide )°D by 1, to get D.
3. Subtract D from each D to get (D - D).
300 << STATISTICS FOR THE SOCIAL SCIENCES

4. Square each (D — D) to get (D — Dy’.


5. Add together each (D — D)’ to get )>(D - D)’.

After finding )* (D — D)’, plug that figure into the S,, formula.

6. Divide }* (D — D)’ by , to get the variance of sample differences.


7. Take the square root of that variance to reveal S,,.

These steps are performed below:

Step 1 Step 3 Step 4


Before After (ke D=-pPs= (D=py=s

5 10 5 2 9
5 6 ip =I 1
7 8 il =] 1
8 7 =k —3 9
i y) ve 0 0
3 i 4 2 3
D131 S12 \\(D-Dy =24 — Step5

Step 2

Mp 6

Steps 6 & 7

F ya 2) eee ae

8. Having established that S,, = 2, we may now complete the ¢ formula.

D 2 2 2/5
Pe = = = anf a= = 256ovat
Soa Oba ose
Thus,

tobtained =2.236 df=n,-1=6-1=5


Two-Sample t Tests 301

Examining the one-tailed ¢criticals at 5 degrees of freedom:

ifcritical Os tevel = 21015:< 2.236 kejectae

Lcxitical 025 level = 2.571 > 2.236 pis OS

We conclude that in the population (p < .05), viewing the commercial


raises favorability ratings for the advertised product.
Many texts also provide alternative computational formulas for the
dependent samples test. Since this procedure is not a widely used test in
social science disciplines where experimentation is not widespread, we will
not cover those computational formulas here.

COMPUTER APPLICATIONS: DEPENDENT SAMPLES ¢ TEST

SPSS

Enter the data, with VAROOOO01 being the six before scores and VARO0002
being the six after scores (see Figure 9.4). Then click

Analyze
Compare Means
Paired-Samples ¢ Test

In the Dialog box on the left, highlight VAROQOO and also VAROOOO2.
Use the arrow button to move them into the Paired Variables box. They will
appear in that box as VAROO001-VARO0002. Click the ok button. The output
appears along with the input in Figure 9.4. The only difference from the
hand-calculated results is that the mean difference and ¢ appear here as neg-
ative numbers. SPSS subtracts the after score (VARO0002) from the before
score (VARO0001), whereas we did the opposite. In any event, the absolute
value of t is 2.236 and df= 5, just as we calculated.
Also, on the printout, the two-tailed probability (Sig.) is given as .076.
Since we were doing a one-tailed test, divide that number by 2 and you get
.038 as the exact probability.

SAS

As before, open the program and click

Solutions
Analysis
Analyst
302 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 9.4 SPSS Data and Printouts for Dependent Samples ¢ Test

Input:

VAROOOO1

5.00

5.00

7.00

8.00

7.00

3.00

Paired Samples Statistics

Std. Std. Error


Mean N Deviation Mean

Pair 1
VAROOOO1 5.8333 6 1.83485
VAROOO02 7.8333 6 1.47196 60093
Paired Samples Correlations

Pair 1 VAROOOO1 6 ;
& VAROOO002

Paired Samples Test

Paired Differences

| 95% Confidence
Interval of
the Difference
Std. Sid. Error
Mean Deviation Mean Lower

Pair 1 VAROOOO1— —2,00000 2.19089 —4,29920


VAROOO02
1

Sig. (2-tailed)
Two-Sample t Tests B® 303

Key in the data: before in columnA and after in column B (see Figure 9.5).
Then click

Statistics
Hypothesis
. Two-Sample Paired ¢ test for Means

In the Dialog box, on the left, highlight A and click the group 1 button;
A appears in the box below the button. Now highlight B, on the left, and
click the group 2 button; B appears in the box below that button. Click
the ok button. The output is presented in Figure 9.5. Again, ¢ is negative
(t = —2.236), df = 5, and you will have to halve the probability of 0.0756
because of the directional alternative hypothesis: p = .0378. (If the negative
t still bothers you, redo the run, putting B in group 7 and A in group 2.)

Excel

Open the program and key in the data: before in column A and after in
column B. Then click

Tools
Data Analysis
t Test: Paired Two Sample for Means

Click the ok button. Highlight the data in column A, moving the infor-
mation into the Variable 1 Range box. Move the cursor to the Variable 2
Range box, highlight column B, and move that information into the box.
Click ok. The output appears on Sheet 4 and is reproduced in Figure 9.6.

STATISTICAL SIGNIFICANCE
VERSUS RESEARCH SIGNIFICANCE

When we initially discussed the one-sample z and ¢ tests, we noted that two
factors have an influence on the magnitude of the statistic generated: the
difference between the two means and the size of the sample.
This can be seen in the original z formula, as algebraically transformed
below.

Zz
_k-n _ &-wi/yn
Tat 0
An increase in either the size of ¥—p or the size of 7 will enlarge the
numerator and thus the final value of z. In the case of the two-sample ¢ test,
304 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 9.5. SAS Data and Printout for Dependent Samples ¢ Test

Input:

A B Cc

1 5 10
2 5 6
3 ih 8
5 8 7
6 7. 9
i 3 7
8

Output:
Two-Sample Paired ¢ test for the Means of A and B

Sample Statistics

Hypothesis Test
Null hypothesis: Mean of (A — B) = 0
Alternative: Mean of (A-— B) “=0

SZ oO 5 0.0756

the same is true except.X,—is replaced by x ,—x,, and the sample sizes may
apply to either 7,, ”,, or both.
Just because the z or¢ is statistically significant and we can reject the null
hypothesis does not mean that the difference between the population
means is large enough to have significance, in the sense of relevance or
importance, to the researcher and reader. To differentiate significance,
Two-Sample t Tests p 305

Figure 9.6 — Excel Data and Printout for Dependent Samples ¢ Test

Input:

A B Cc

1 5 10
Z 5 6
2s) iy 8
>) 8 7
6 yf 2
v 2) vy
8

Output:
t Test: Paired Two Sample for Means

Variable 1

Mean 5.995539 7.833333

Variance | 3.366667 2.166667

Observatio 6 6

Pearson C 0.135761

Hypothesiz 0

df 5)

t Stat —2.236068

|pere=t) on 0.037793 |

t Critical or 2.015048

P(T<+t) tw 0.075587

ltCritical tw 2.570582
306 STATISTICS FOR THE SOCIAL SCIENCES

meaning relevance, from statistical significance, we call the former meaning


research significance. That is what we mean when we use the term signifi-
cance in everyday life. Thus, the two meanings of significance are defined
as follows:

Research Significance: Relevance; importance of a particular differ-


ence of means or other finding.

Statistical Significance: The high probability that the difference


between two means or other finding based on a random sample is not
the result of sampling error but reflects the characteristics of the popu-
lation from which the sample was drawn.

Whether or not a finding that is statistically significant also has research


significance is solely up to the judgment of the individual! The general
public—and too many researchers—fail to differentiate statistical signifi-
cance from research significance: relevance. Even if the difference between
two means is trivial, if 7 is large enough, that difference will be statistically
significant! Be particularly wary of statistically significant results gleaned
from large data sets. Research significance is in the eye of the beholder!
Be aware also that the term significant can be used to mislead the
layperson who hears that there is “a significant difference between deter-
gent A and detergent B,” or “our group got significantly fewer colds by
taking Vitamin C,” or any context in which the term significance, meant
to mean statistical significance, will automatically be interpreted to mean
research significance or relevance. The latter conclusion must be based not
on the size of the calculated z or ¢ but on whether the difference between
the two means seems significant to the observer.

STATISTICAL POWER

An additional factor affecting our ability to reject the null hypothesis is the
nature of the selected test of significance itself. This brings us to the concept
of statistical power, the likelihood that our test will reject the null hypoth-
esis when, in fact, H, really is true. How likely is our test to reject a null
hypothesis when the null hypothesis is false and “ought to be” rejected?
sarlan
i an
cr aas eh
Statistical power =The likelihood that our test will reject the null hypothesis when,
in fact, H, really is true.
MMOL
N MAMAN NN ONCE ENN
n
Two-Sample t Tests j» 307

Up to this point, we have been concerned about Type I or alpha error,


the probability of mistakenly rejecting a true null hypothesis. However
,
there is another type of error, a Type II error or beta error, which
exists
but is not reported. Beta is the probability that the null hypothesis is really
false—H, is true—but our obtained statistic—z, 4 and so on—was too
low to enable us to reject the-H,, even though it “ought to be” rejected. The
relationship between these two types of errors is illustrated in Table 9.4.

sss

Type Il error or beta error The probability that the null hypothesis is really
false—H, is true—but our obtained statistic—z, t, and so on—was too low to
enable us to reject the H,, even though it “ought to be” rejected.
SOMES
EER RS ARN HN

The probability of each occurrence is listed in the appropriate cell of the


table. If H, is really true and we conclude H, from our test, we make a Type
I error whose probability is alpha. If H, is really false but we fail to conclude
it with our test, we make a Type II error whose probability is beta. What if we
reach the correct conclusion? Suppose H, is true and our test correctly fails
to reject it. Since the sum of all possible probabilities is 1 and alpha is the
probability of falsely rejecting H,, the probability of correctly not rejecting lek
is 1 — a. Similarly, if H, is true and we conclude that fact, the probability of
that occurrence is 1 — B. Notice, therefore, that statistical power as we have
defined it is really one minus beta! Power = 1 — B.
We must point out that alpha and beta are related in the sense that by
setting alpha low, as we traditionally do, we increase the probability of beta.
By making it hard to reject the null hypothesis, we also increase the likeli-
hood of failing to reject null hypotheses that “ought to be” rejected since
they are in fact false.
Thus, there is a clash of interests between the traditional conser-
vatism ofstatistical tests (which require a low alpha and thus a large beta)

Table 9.4

Based on Our Significance Test, We Conclude:

In the Population: H, Is True H, Is True

Hi, is true No error Type | error


p=1-oa p=a
H, is true Type II error No error
p=B je eal
308 STATISTICS FOR THE SOCIAL SCIENCES

and the objectives of social and behavioral researchers who are seeking
to identify the relationships between variables or the real differences
between groups in the population. Such differences must be rather large
if we are to reject the null hypothesis with our tests of significance.
Consequently, many relationships between variables or differences
between population means that really do exist are not large enough to
yield statistically significant results.
For example, in evaluation research, Lipsey (1990; see Note 2) found
that only 28% of the time do small effects (a hypothesized .20 difference
between yu for the experimental group and yp for the control groups) yield
statistical significance. In 72% of the studies, no significant results would be
detected. Studies of research in other social sciences indicate that between
18% and 34% of the time, such small mean differences yield significant
results. An exception was sociology at 55%. (Due possibly to larger available
sample sizes?) For medium effects (a hypothesized .5o difference between
the population means), the percentage of studies yielding significant results
was between 52% and 76% of the time, depending on the discipline (with
sociology again being higher at 84%). Large effects (a .80 population
mean difference) yielded significant results between 71% and 94% of the
time for most social sciences. Clearly, only larger mean differences are found
to be statistically significant a majority of the time in social research.’

oases

Small effects Hypothesized .2o difference between yu for the experimental group
and wu for the control groups.

Medium effects Hypothesized .50 difference between the population means.

Large effects A .80 population mean difference.


seo peste

Cohen (1977, 1988; see Note 2) has suggested that a reasonable beta
value be set at .20. Thus, statistical power (1 — B) should be .80 at a mini-
mum. Lipsey (1990) identified four factors that determine statistical power:
the test itself, the alpha level, the sample size, and the effect size as esti-
mated (for the two-sample ¢) by the difference between the two sample
means or from other sources. It turns out that of the four factors, the test to
be used is determined largely by the type of data available, and alpha is
determined by statistical tradition. Thus, the sample size and effect size
remain just as we demonstrated earlier by the z formula. Of these two, it can
be demonstrated that the effect size has a far larger impact on statistical
power than an increase in sample size. Tables have been developed showing
the relationship between statistical power, sample size, and effect sizes for
Two-Sample t Tests j» 309

differing alpha levels. Consult the two works cited in this section for more
details on the use of such tables (see Note a
A word of caution: Statistical power and effect size appear to be of most
concern in the social sciences that must make use of rather small samples
or experimental and control groups. Where samples can be larger, such as
in sociology, the problems are diminished. You will have to inquire in your
own discipline as to the current level of concern about power and effect
size. While the low level of statistically significant research results is a matter
of concern to us all, to a traditionalist, much of this concern may appear
to be a justification for the use of levels of significance Jower than the ones
traditionally applied. If so, it will remain controversial.

CONCLUSION
The two-sample ¢ test is one of the oldest tests of statistical significance
and one of the most commonly encountered tests. Since comparison of two
groups’ means is a common method of data analysis, and often the groups
being compared are random samples, there are many situations that make
use of this test. Moreover, the two-sample ¢ test may be used in both
experimental and nonexperimental situations, as the examples presented in
this chapter demonstrate. In fact, you will find it used in just about every
discipline employing statistical techniques.
Now that we have studied the ¢ tests for the difference between the means
of two samples (or randomly assigned groups), we can turn to the problem of
comparing a larger number of sample means, using the technique of analysis
of variance.

Chapter 9: Summary of Major Formulas


ht rms |

q Independently Drawn Samples


-—_———
F Test for Homogeneity of Variables

2 df numerator = 7 — 1 for the sample with the larger


i sale variance
Ssmaller df denominator = 7 — 1 for the sample with the smaller
variance
hee es
(Continued)
310 << STATISTICS FOR THE SOCIAL SCIENCES

(Continued)

The Two-Sample t Test Calculated From S$

Equal Population Variances Assumed


x
t= are af =m +nz=2
misj tmp
ni+n2—2
(1
ny

Unequal Population Variances Assumed


x1 —X2 df Estimate: whichever 7 is smaller

sy 3 df Exact: .
d = 2 2 or

ie ee

[SS
The Two-Sample ¢ Test Calculated From 6?

Equal Population Variances Assumed


X1 —X2
CS SSS eae 2
(n\—1)67 +(2—-1)aF Ae as22
ny+n2—2 ny n2

Unequal Population Variances Assumed

x1 — X2 df Estimate: whichever 7 is smaller


pS
or : a
df Exact: (3
ny
rs 2)
n2
df = 5
(=)
(1,—1) a8

Dependent Samples

— gite Lao where Sp = delDiralde


Sp/fMp — 1 a Np
Two-Sample t Tests » 311

Exercise 9.1 _
A tolerance index has been developed that is designed to measure one’s tolerance of
“unpopular” beliefs such as those of a racist or sexist nature. On the scale, 0 means
the lowest level of tolerance and 15 the highest level. A random sample of 10 uni-
versity students (Group 1) is scored along the index. A second sample (Group 2) of
students from the same university is a sample of students who had recently attended
-a workshop on multicultural diversity. Making no directionality assumption in H,, test
3 pullhypothesis that there is no difference in tolerance between the two populations.
Group 1 Group2
(Control) (Workshop)
x, x=

DouUAnbRWWH si
Gra
kB
oo
AP

Exercise os.
Refer tot Exercise 9.1. Suppose you had prior evidence that people attending such
workshops generally demonstrated increased tolerance. What are your conclusions
_ with a directional He

Exercise 9.3
The control group of students from peu 9.1 is compared to a random sample
ot military veterans attending the same institution. Making no. directionality
assumption, test for significance.
: : Group 2
(Veterans)
x=

MOOONnNuUoRAR
312 STATISTICS FOR THE SOCIAL SCIENCES

Exercise 9.4
In the SAS® printout below, the control group from Exercise 9.1 is compared to a
sample of fine arts majors at the same university. Select the appropriate t test and
state your conclusions about the two populations. (The standard deviations use
formulas with n — 1 in the denominators.)
SAS
TITEST. PROCEDURE
VARIABLE: SCORE
GROUP N MEAN STD DEV STD ERROR MINIMUM
1 10 4.30000000 1.33749351 0.42295258 2.00000000
2 10 ZoO0000000. 2.63523138 0.83333333 5.00000000

MAXIMUM VARIANCES 1s DF PROB > |T|


6.00000000 = UNEQUAL —3.4242 13.3 0.004
10.00000000 EQUAL —3.4242 18.0 0.003

FOR HO: VARIANCES ARE EQUAL, F’ = 4.88, df = (9.9) PROB > F’ = 0.0559

Exercise 9.5
The control group from Exercise 9.1 is now compared to a random sample of the
faculty from the liberal arts college. Test for significance.
Group 2
(Liberal Arts Faculty)
AY,

Exercise 9.6
Below is the SAS’ printout for Exercise 9.5. What is the difference between this and
your findings? Does this change your conclusion about significance?

VARIABLE: SCORE
GROUP N MEAN STD DEV VARIANCES T DF PROBS 1H
| 10 4.30000000 1.33749351 UNEQUAL ~—3.3149 10.2 0.0077
2 10 10.00000000 5.27046277 EQUAL 3.3149 18.0 0.0039
FOR HO: VARIANCES ARE EQUAL, F’ = 15.53 DF = (9.9) PROB > F’ = 0.0004
Two-Sample t Tests ® 313
314 <4 STATISTICS FOR THE SOCIAL SCIENCES

Faculty
Before After

5 10
5 10
5 5
5 0
5 0
15 10
15 10
15 15
15 15
15 3

Exercise 9.11
Set up and run on the computer the following:
A. The data from Exercise 9.1. Do your results resemble the ones that you
calculated earlier?
B. Replace Group 2 with the veterans’ data in Exercise 9.3 and compare the
computer results to those you previously calculated. Do they coincide?

NOTES
1. If the df we need does not appear in the table, we use the adjacent df value
that makes it harder to reject H,. For example, if 7, had been 10 instead of 4, we
would have selected either 8 or 12. F..,.,.,, where 7, = 5 is 4.82 where 7, = 8, and 4.68
where 7, = 12. Since 4.82 is the larger value, we use it as F.,.,-.-
2. The seminal work in this area is to be found in J. Cohen, Statistical Power
Analysis for the Behavioral Sciences, 2nd ed. (Hillsdale, NJ: Lawrence Erlbaum,
1988). However, M. W. Lipsey, Design Sensitivity: Statistical Power for Experimental
Research (Thousand Oaks, CA: Sage, 1990), is much easier reading for the beginner.
It was Cohen who operationalized large effect sizes as .8, medium as .5, and small as
.2, These effect sizes are projected population mean differences expressed (as in the
z test) in standard deviation units.
3. These tables are from a previously issued version of SAS.
~ > =
- = —— a -
~ se
: s

WAP Ttr 10 -

. - >
IO
ee :

ie W +a Amalys is
S23
ass
ene |! ae =

-_

© Meig>os m4
| re , Wt mays ae QP »
Pan ae
WTP \eapMii » aay ioBese ©
at ie
OY

eee | ee et : Sie, ine Omir


'

: at >-s-Onpeinrt, - - ares Ve ve id

AP
a)

hd
_ i wee
;
ee
; _
=
4
fw Sil ss :
wa
7amby oee
i> @& s Ss)
re — et -
- « - .
eet)
7 j (yeeq eae
ys oo a 18 Y a iy Cet

F a 6 _
Cita, fl ay

spef: orenthySire eau 54


al a
4 Fheed ial 7 Prine a)
: ’ 7 7
9 —_ ss a
- 7
eae or v
a a
an 7
YW KEY CONCEPTS V¥
EB OLSSON RO Pe NE ES EE TORN IIR SCE EAR GONE OES LELEE NONOE ELD TT DALE EN RE SE

one-way analysis of within-groups robustness (of a test


variance/one-way sum of squares/ of signifcance)
ANOVA SSwithin/'SSwy post hoc procedures/
F/the F ratio error sum of squares post hoc test of
participants or subjects between-groups mean multiple comparisons
(in an experiment) square/MS),weer MS, Scheffe’s S test
grand mean degrees of freedom, Scheffeé’s S critical value
sum of squares/SS between-groups/ MODEL SS
mean square/MS A screen! MODEL MS
total sum of squares/ within-groups mean ERROR SS
SSrar/ SS square/MS\,.i/MSw,
ithin two-way analysis of
between-groups degrees of freedom, variance/two-way
sum of squares/ within-groups/df yirnin/Aw ANOVA
SS Between hoe ANOVA source table interaction effects
RELL ELE IEELNELILE ANE LEEELEERELL LD LE LD ELLE LEE LEE LEN LN EEOOASSEN EBC RE BES ES EN ES lar NN nc og
CHAPTER

One-Way Analysis
of Variance

VY PROLOGUE ¥

This chapter expands the kinds of comparisons in the last chapter to more
than two groups as well. So this time, our juvenile criminals might be
broken into three groups: one of those that had detention, one of those
with probation alone, and one group that received probation and com-
munity service. In our marriage counseling example from before, suppose
our couples are assigned randomly to three groups: individual counsel-
ing, group counseling (with more than One couple participating), and no
counseling.
We are not limited to three groups. Maybe we are doing a study for
a pharmaceutical company. We randomly assign participants with the same
illness to five groups as follows: one getting ¥2 mg per day, one getting 1 mg
per day, one getting 1/2 mg per day, one getting a placebo (a “sugar pill” with
no therapeutic value), and one group getting no medication at all. We then
compare the mean recovery time for each of the groups.
gases

p> 317
318 << STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
Analysis of variance (ANOVA) is a statistical cousin to the ¢ test. Like the
t test, it is a technique for comparing sample means, but unlike the ¢ test,
ANOVA can be used to compare more than two means. Analysis of variance
is very versatile. It is particularly friendly to experimental applications, where
we may be comparing the means of several treatment groups and a control
group. Consequently, psychologists rely heavily on this procedure. ANOVA is
also useful in nonexperimental situations in the same way that the f test is.
Interestingly, though, ANOVA has been less widespread in nonexperimental
research than many other statistical procedures, despite its great potential.

One-way analysis of variance (ANOVA) A technique for comparing sample means


that can be used to compare more than two means.

With ANOVA, because several sample means are usually being compared,
once a null hypothesis has been rejected, we need a follow-on, or post hoc,
procedure. This is because although ANOVA examines all sample means at
once, it is possible that some pairs of means may not be significantly differ-
ent from one another, even though when all means are taken together in
their entirety, the null hypothesis may be rejected. Thus, the process is a bit
like remote sensing (i.e., aerial photography). ANOVA gives us a high-altitude
picture, and if we can reject the null hypothesis, we swoop down for a closer
look. The post hoc test provides the low-altitude shot.
essen blo
ESEHH HOMESS

Post hoc procedure A follow-on procedure that is used once a null hypothesis
has been rejected.
i eiaiiaaeaeiileeaediaaneeaeaeeea amen adden mmr commer ee eee

At the end of the chapter, we will briefly look at some other variations
of the ANOVA technique.

HOW ANALYSIS OF VARIANCE IS USED


Analysis of variance is designed, in its nonexperimental application, for
two variables; one is interval level of measurement (usually the dependent
variable), and the other is grouped data of any level of measurement.
Suppose we study a random sample of 8 people. We determine for each of
them whether they are from urban or rural areas as well as their score on a
“pro-life” index that ranges from 0 (most pro-choice) to 200 (most pro-life)
on the issue of abortion. We get the following results:
One-Way Analysis of Variance » 319

Rural (Group 1) Urban (Group 2)

160 100
130 110
150 130
iN Ole 120
y
\x, = 580 yx, = 460

us 580 = 460
Mie ag ibs Bo SirNeer stl

Clearly, the means of the two groups differ in our sample of 8 subjects.
May we assume that they differ in the population as a whole? Assuming no
directionality, our hypotheses would be as follows:

Fy: [, =

FL,: LL, # UM,

We could handle this problem using the two-sample ¢ test, or we could


perform a one-way analysis of variance (one-way ANOVA). Unfortu-
nately, analysis of variance does not allow for a directionality assumption
(although ANOVA’s cousin, the two-sample ¢ test, does). Thus, we use the
nondirectional H,.

One-way analysis of variance (one-way ANOVA) Analysis of variance with one


dependent variable and one independent variable.

From our sample data, we will calculate a statistic called & or the F ratio
(named for Fisher, who originally helped developed it). As we did with other
tests, we will compare our obtained F to F critical at the .05 level, and if our
F exceeds F ica) We Will reject the null hypothesis.

For Fratio The statistic generated by the analysis of variance procedure.

ANALYSIS OF VARIANCE IN EXPERIMENTAL SITUATIONS

Our first example (abortion stance by area) came from a nonexperimental


context, a survey. In many social sciences, experiments have been relatively
rare, but in the behavioral sciences, laboratory sciences, and medicine, they
320 @ STATISTICS FOR THE SOCIAL SCIENCES

are quite common. Where experimental designs are common, ANOVA is the
most common statistical technique used.
Let us use an example from training and development, a growing
subfield of communications that deals with the training of adults, usually
in a job-related setting. Suppose an advertising firm is seeking to improve
its operation. A training program is being established for employees of this
firm to give the employees skills needed to properly advise and assist
clients. The firm wishes to develop the most effective training program
possible. The trainers are interested in comparing the relative efficiency
of day, night, and weekend programs.
Let us assume that the same instructor will present identical material in
each of three sections. One section will meet 1 hour per day, Monday through
Friday, at the same time for one week. The second section will meet at night
under similar circumstances. The third section will meet on a Saturday for a
single long session, including 5 hours of instruction plus time for breaks and
meals. There will be 10 people in each section. At the end of the instruction,
each person taking the class will rate his or her overall satisfaction with the
course on a scale of 0 (dissatisfied) through 5 (completely satisfied).
The 30 “students” are the participants or subjects, as they used to be
called, in the experiment. They are all employees of the same firm and are
selected to take the course for professional purposes. There is no random
sample being selected here. However, if the trainer suspects that satisfaction
will differ among the three classes—day, night, and Saturday—he or she may
test that hypothesis.

Participants or subjects The people being studied in an experiment.

To do so, a table of random numbers or some computerized random


number generator can be used to randomly assign the 30 subjects to three
groups of 10 people each. This randomization process tends to control for
other social or behavioral characteristics. Each group will, for instance, have
about the same proportion of men and women as the original group of 30.
Each group will have about the same mean age. Each group will have about
the same proportion of disgruntled employees who feel forced by their
bosses to take the course. (We keep saying “about the same” since there will
be some differences, slight we hope, equivalent to sampling error, resulting
from chance in the randomization process.) Sampling error aside, if we ran-
domly assign subjects, any difference in mean satisfaction scores among the
three groups will be the result of the time frame for the course (day, night,
Saturday) rather than other factors. Thus, for the variable satisfaction:

AA): Hany ae Haight ar Ussturday


One-Way Analysis of Variance j» 321

Note that analysis of variance will take as many groups (or categories)
as we have, whereas the two-sample ¢ test is limited to two groups (or
categories). Thus, for the above problem, we are only able to use one-way
analysis of variance.
With more than two categories, it becomes difficult to write H, with sym-
bols. In effect, H, says that in the population, there exists at least one
inequality that negates the null hypothesis. Any of the following would
negate H).

Maas a Piiete

Haay cesaturday

Haight = saturday

It is 20t necessary that all three population means be unequal, although that
could be the case:

Haay = Haight eosaturday

Suppose we obtain the following results:

x = Satisfaction katings

Day Night Saturday

nee
oper
er
Pete
Ge
6
4 3 S \| iN pa
ed os \| Oe
ARR
oes
onnou
On
Nes
i) eS x Il Oe
SINWWW
KUUY
AA

=4.1 Kv — =2.6 Cae Be) 6)


Xp =—

from each
Clearly, the category means of 4.1, 2.6, and 3.8 are different
)

de that they were not the result


other. Are they different enough to conclu
n proces s whereby
of sampling error—in this case, due to the randomizatio
322 << STATISTICS FOR THE SOCIAL SCIENCES

we assigned the subjects to the three categories? If we can reject the null
hypothesis, we may conclude that “in the population,” that is, for people in
general, student satisfaction levels differ by the time and format of the class
offered regardless of the instructor, course content, or anything else.
We calculate F and compare it to F.,,,,.4) at the appropriate degrees of
freedom. If the F we obtained exceeds [Link]? .05 level, we reject H,. We then
compare our F to F.,,,,.,, at other levels to form a probability statement.

F: AN INTUITIVE APPROACH

We may think of F as a measure of how well the categories of the indepen-


dent variable explain the variation in the scores of the dependent variable.
If the categories of the independent variable are totally useless in explaining
the variation of scores of the dependent variable, then F will equal zero. As
the categories of the independent variable begin to explain or account for
some of the variation in the dependent variable, F begins to get larger. The
better the independent variable explains variation in the dependent vari-
able, the greater is the relationship between the variables and the larger F
becomes. When the categories of the independent variable explain nearly all
the variation in the dependent variable, F grows extremely large, approach-
ing infinity as a limit.
As an illustration, imagine a simplified version of the course satisfaction
problem. For simplicity’s sake, we will limit the independent variable to two
categories, a day class and a night class. Suppose there are six people in each
class. The following satisfaction scores emerge.

Case 1

oS = = si= ™

5 5 In this case, F—which you


5 5 haven't yet learned to
5 5 calculate—would be zero.
S 3 F=0
3 3
3 2 Grand Mean = 48/12 = 4.00
One-Way Analysis of Variance » 323

If we were to ignore the existence of the two categories and just calculate
the mean of all 12 satisfaction scores, the mean we would get—called the
grand mean—would be 48/12 or 4.0, the same as the two category means.
Since mean satisfaction is the same (4.0) whether or not we know in which
category a subject belongs, the categories do not help us predict a subject’s
score on the dependent variable. Thus, the two variables are unrelated. There
is no difference between day and night classes in terms of course satisfaction.
F will be zero.

Grand mean The mean of all scores.

Now imagine a slight variation in which one more person in the day
class has a satisfaction score of 5 and one more person in the night class has
a satisfaction score of 3. While the grand mean is unchanged (4.0), the two
category means now differ.

Case 2

Day Night
5 5 Pie As
5 5
3) 4 Grand Mean = 48/12 = 4.00
5 3
5 a
3 3
S20 2

2,
ip = > = 433 Xn = — = 3.67

Once we learn to calculate F we will see that for this problem, F = 1.248,
up from 0 in Case 1. Whereas in Case 1, each category had three scores of 5
and three scores of 3, in Case 2, the day class is slightly more satisfied than
the night class: four scores of 5, two scores of 3 in the day class with a mean
of 4.33 as opposed to two scores of 5, and four scores of3 in the night class
with a lower mean of 3.67. The two classes now differ somewhat in terms of
satisfaction.
We now add one more score of 5 to the day class, replacing a score of 3,
and in the night class, we replace a score of 5 with a 3.
324 4 STATISTICS FOR THE SOCIAL SCIENCES

Case 3

Day Night
5 5 F=7.994
> 6:
5 8 Grand Mean = 48/12 = 4.00
5 5
5 3

ia 2
3 y= 20
e
= 28 9 20
== 2 7 go earn eo

The differentiation between day and night classes, in terms of satisfaction,


has grown even more pronounced. Day students are clearly more satisfied
than night students. The category means are more widely spread about the
grand mean than in Case 2. Though we could be wrong, there is now a five-
to-one chance that if someone tells us he or she is in the day class, we would
be accurate in predicting the satisfaction score as being 5. (Of the six day
students, five have scores of 5; only one does not.) If someone is in the night
class, we predict (also with odds of five to one) that the satisfaction score is 3.
Finally, we change the final score of 3 in the day class to 5 and change
the final 5 in the night class to 3.

Case 4

Day Night
5 3 F is mathematically undefined,
5 but had been getting larger,
5 3 approaching infinity as a limit.
5 3
5 3 Grand Mean = 48/12 = 4.00
2 3
» ey re erga ls

a 30 mn 18
to ae BS ee

In this case, the categories of class type explain all the differences in
satisfaction scores. Knowing what class one is in gives us perfect predic-
tive ability in terms of satisfaction. If in the day class, a person’s satisfaction
One-Way Analysis of Variance 325

score is 5; if in the night class, the satisfaction score is 3. The two variables,
satisfaction and class meeting time, are perfectly related.
Notice that in Case 1, all the variations of the scores from the grand mean
were actually within each of the two categories. The category means did not
vary at all from the grand mean. As we progressed through Cases 2 and 3,
more and more of the deviatioris or variations of scores from the grand mean
could be explained by the category means. Finally, in Case 4, there were no
deviations of scores within the categories. All scores fell at the means of their
respective categories. All deviations of scores from the grand mean could be
accounted for by the deviations of their respective category means about the
grand mean. To see this more clearly, note that algebraically, we can break
the distance between any score and the grand mean into two components:
(a) the distance from that score to its respective category mean, plus (b) the
distance from that category mean to the grand mean.

(Score — Grand Mean) = (Score — Category Mean)


+ (Category Mean — Grand Mean)

In Case 1, every score is either a 5 or a 3, both category means are 4.00, and
the grand mean is 4.00. Thus, for a score of 5 in the day class,

(> —A) = (5-4) + (4-4)

The alc] The Jal The Grand Mean


The Grand Mean The Category Mean

(5-4) =(5-4)+ (4-4) This part reduces to zero.


(1) = (1) + (0) |
=sa Qx:

The same would hold for a score of 5 in the night class. For a score of 3 in
either class,

(3 -4)=(-4)+ (4-4) This part reduces to zero.

reduces to
Notice that in Case 4, it is (Score — Category Mean) that always
zero. Since all scores in the day section are 5,
@ STATISTICS FOR THE SOCIAL SCIENCES

(5-4) = (5 =5) ts (5 =4)

The al The ne The Grand Mean


The Grand Mean The Category Mean
v
a

(5 - 4) =(5—5)
+(5-4)
This part reduces to zero.

Since all scores in the night section are 3,

(3 —4) = (3.= 3) 35 (3
— 4)

The tel The ty The Grand Mean


The Grand Mean The Category Mean

(33 -4)=(3 -3)+G3 -4)

This part reduces to zero.

In Case 1, (Category Mean — Grand Mean) is always zero; in Case 4,


(Score — Category Mean) is always zero. These are the two extreme cases.
In Case 1, the categories explain no variation in the dependent variable, and
in Case 4, the categories explain all the variation in the dependent variable.
Cases 2 and 3 fall in the middle. In Case 2, one third of the distance between
a score and the grand mean can be accounted for by the distance between
the category mean and the grand mean; two thirds of the distance is within
the category, between the score and its category mean. In Case 3, two thirds
of the distance between a score and the grand mean can be accounted for by
the distance between the category mean and the grand mean; one third of
the distance is within the category, between the score and its category mean.

ANOVA TERMINOLOGY

The logic of analysis of variance is based on partitioning the distance to the


grand mean into distances explained by the category means and distances
unexplained by the category means, in a manner somewhat analogous to
One-Way Analysis of Variance jp» 327

what we have been doing. However, ANOVA squares distances to get rid
of negative numbers, and it works with sums of these squared distances.
It also makes use of its own computational formulas. Thus, before learning
the technique for calculating F we need to define some terminology.
Since we will ultimately be using variance estimates (squared sigma-
hats), let us return fora moment to the definitional formula for a population
variance estimate.

pe aie aa)
o- =
n—1

The denominator for this estimate, 7 — 1, is its degrees of freedom


or df The numerator SS (x—X)* can be read as the sum of the squared
deviations of the values of x from the mean. Each x—x is the deviation,
the distance of that value of x from the mean. These deviations are
squared to get rid of negative signs, and the squared deviations are
summed. In analysis of variance, “the sum of the squared deviations of
the values of x from the mean” is shortened to the sum of squares and
is indicated by the letters SS.

Sum of squares The sum of the squared deviations of the values of x from the mean.

When we divide SS by ” for a sample variance or by dfn = 1 for a vari-


ance estimate, we are finding a kind of average amount of squared deviation
for each value of x: “the mean squared deviation of a score (a value
of x) from the mean of all scores,” which is shortened to the mean square
and is indicated by the letters MS. A variance or variance estimate is the aver-
age, or mean, amount of the sum of squares per unit. Our 6” formula is then
symbolized as

Mean square The mean squared deviation of a score (a value of x) from the
mean of all scores.

What ANOVA does is first to calculate, using the computational formula,


the total sum of squares (SS,,,,, or SS,), meaning the total of the squared
deviations of scores about the grand mean.
328 < STATISTICS FOR THE SOCIAL SCIENCES

Total sum of squares The total of the squared deviations of scores about the
grand mean.

io = pa a (Sa)"
1

The total sum of squares (SS,) is then partitioned (divided) into two
components. The first component is the between-groups sum of
squares (SS, ...ce, OF SSg), the portion of the total sum of squares that can be
accounted for by the variations of the category means about the grand
mean. That is, SS, is the portion of SS, that can be accounted for (explained
by) the categories.

Between-groups sum of squares The portion of the total sum of squares that can
be accounted for by the variations of the category means about the grand mean.

The second component is the within-groups sum of squares


(SS\iehin OF SSy), the portion of the total sum of squares left unexplained by
the variations of the category means about the grand mean. This, then, is
the sum of squares within the categories or the squared deviations of
scores about their respective category means. It is sometimes called the
error sum of squares.

Within-groups sum of squares or error sum of squares The portion of the total
sum of squares left unexplained by the variations of the category means about
the grand mean.
smeaeeeeeeemaeaneaeeeeneemeenenemeneeeacenememeereeeeemmeieameememmeninememmmnmmmmnenn
mma smmememmmmemmmmmmmrirm er

In short,

SS;, = the portion of SS, accounted for by the categories of the indepen-
dent variable.

SSy = the portion of SS, not accounted for by the categories of the inde-
pendent variable.

These are additive:

SS, So, Sse,


One-Way Analysis of Variance jp» 329

We use SS, and SSy, to form two separate population variance estimates.
The first of these variance estimates, the between-groups mean square
(MS perween OF MS3), is a variance estimate based on the between-groups sum
of squares. MS, estimates the population variance accounted for by the vari-
ation of the category means about the grand mean—the population vari-
ance accounted for by the groups or categories of the independent variable.
To find MS,, we divide SS, by the between-groups degrees of freedom
(Af, OF Afserween): Since here we are talking not about the number of respon-
dents but about the number of groups or categories, df, equals the number
of categories (or groups) minus 1.

Between-groups mean square A variance estimate based on the between-groups


sum of squares.

Between-groups degrees of freedom The degrees of freedom based on the


number of groups studied or categories of the independent variable.

af, = no. of categories — 1

SO

SSp SSp
MSR =
df, no. of categories — 1

The second variance estimate is the within-groups mean square


(MSyin.,,
ithin
OF MSy,), 4 population variance estimate based on what the cate-
gories of the independent variable do not explain—the variation of scores
within the groups. We take SS,, and divide by the within-groups degrees
of freedom (df,,,,,. OF Gy). The within-groups degrees of freedom is
found by subtracting the between-groups degrees of freedom from the
degrees of freedom belonging to 6* (our original total variance estimate),
which is 2 — 1. Thus,

Within-groups mean square A population variance estimate based on what the


categories of the independent variable do not explain—the variation of scores within
the groups.

Within-groups degrees of freedom That portion of the total degrees of freedom not
accounted for by the number of groups studied.
Sree SeSN IRONS SESE LESSSS ETE NT ET ETS SEES,
330 << STATISTICS FOR THE SOCIAL SCIENCES

fy = Ara — Upermeen = (1 — 1) — (no. of categories) — 1

sO

eee SSw SSw


oe a Af w, ~ (n—1) — (no. of categories — 1)

Note that like the sums of squares, the degrees of freedom are additive, so that

Af pat = dfs a df

af, + afy = (no. of categories — 1)


+ [(n — 1) — (no. of categories — 1)]
Sa i Af ora

Although the SSs and dfs are additive, the MSs are not. The mean square
total, 6’, does not equal MS, + MSy.
Finally, we find the F ratio:

a MS,
a MSw

The F that we obtain is compared to F.,,,i-, (See Table 10.1).' Note that F
uses two degrees of freedom, [Link], Which is found in the column on the
left-hand side, and df,,,,,,,,, which we locate on the top row. (Here, 7, and 7,
mean degrees of freedom: 7, = df, and n, = df,.) There are three pages to
this table: one for the .05 level, one for the .01 level, and one for the .001
level. If Forineqd CEXCCCAS F viricg at the .05 level, we reject the null hypothesis
and20 On tothe page with.F. at the.00 level. ity exceeds’ 40. at
the .01 level, we go on to compare it to F,.4, at the .001 level. This is exactly
what we did with earlier tests. Probabilities are reported the same way.
If ANOVA is done on a computer using a program such as SAS or SPSS,
the exact probability will be listed, and thus it will not be necessary to use
a critical value of the F table.

THE ANOVA PROCEDURE

Before we actually work an F problem through, look at Box 10.1, where the
computational steps and all appropriate formulas are given. To calculate &
we must find the following: 72 for each category, jy, _X for each category,
Mou and Dx;.,.. which we get by finding }°x* for each category and
adding them up. Note that all the examples used so far in this chapter have
One-Way Analysis of Variance » 331

Table 10.1 — Critical Values of F forp = 05

n\n, 1 2 5) 4 5 6 8 12 24 oo
1 161.40 199.50 215.70 224.60 230.20 234.00 238.90 243.90 249.00 254.30
2 18.51 19.00 19.16 = 19.25 19.30 19.33 19.37 19.41 19.45 19.50
3 10.13 9.55 928° 9:12 9.01 8.94 8.84 8.74 8.64 8.53
4 ial 6.94 6.59 6.39 6.26 6.16 6.04 5.91 5.77 5.63
5 6.61 5.79 5.41 5.19 5.05 4.95 4.82 4.68 4.53 4.36
6 5.99 5.14 4.76 4.53 4.39 4.28 4.15 4.00 3.84 3.67
7. 5) 4.74 4.35 4.12 2 97 3.87 3.73 3.57 3.4] 3.23
8 5.32 4.46 4.07 3.84 3.69 3.58 3.44 3.28 3.12 2.93
9 S12 4.26 3.86 3.63 3.48 3.37 3.23 3.07 2.90 OVAL
10 4.96 4.10 3.71 3.48 3.33 3.22 3.07 2.91 2.74 2.54
11 4.84 3.98 3.59 3.36 3.20 3.09 2.95 2.79 2.61 2.40
12 4.75 3.88 3.49 26 3.11 3.00 2.85 2.69 2.50 2.30
13 4.67 3.80 3.41 3.18 3.02 2.92 207 2.60 2.42 2.21
14 4.60 3.74 3.34 3.11 2.96 2.85 2.70 2.53 235 2.13
15 4.54 3.68 3.29 3.06 2.90 2.79 2.64 2.48 2.29 2.07
16 4.49 3.63 3.24 3.01 2.85 2.74 2.59 2.42 2.24 2.01
ty 4.45 3.59 3.20 2.96 2.81 2.70 255 2.38 2.19 1.96
18 4.41 3.55 3.16 2.93 277 2.66 2.51 2.34 2.15 1.92
19 4,38 B52 3.13 2.90 2.74 2.63 2.48 2.31 aid 1.88
20 4.35 3.49 3.10 2.87, 27 2.60 2.45 2.28 2.08 1.84
pail 4.32 3.47 3.07 2.84 2.68 257, 2.42 2.25 2.05 1.81
22 4.30 3.44 3.05 2.82 2.66 2.55 2.40 2.23 2.03 1.78
23 4.28 3.42 3.03 2.80 2.64 2.53 2.38 2:20. 2.00 1.76
24 4.26 3.40 3.01 2.78 2.62 25 2.36 218 1.98 1.73
25 4.24 3.38 2.99 2.76 2.60 2.49 2.34 2.16 1.96 17
26 4.22 3.37 2.98 2.74 2.59 2.47 2.32 D5 1.95 1.69
OF 4.21 3.35 2.96 2.73 251 2.46 2.30 2.13 1.93 1.67
28 4.20 3.34 2.95 27h 2.56 2.44 2.29 212 1.91 1.65
29 4.18 309 2.93 2.70 2.54 2.43, 2.28 2.10 1.90 1.64
30 4.17 3.32 2.92 2.69 2.53 2.42 227 2.09 1.89 1.62
40 4.08 3.23 2.84 2.61 2.45 2.34 2.18 2.00 1.79 151
60 4.00 3.15 2.76 252 237, 225 2.10 1.92 1.70 1.39
120 3.92 3.07 2.68 2.45 2.29 27 2.02 1.83 1.61 1.25
co 3.84 2.99 2.60 2.37 2.21 2.09 1.94 1.75 152 1.00

Table 10.1 Continued—Critical Values of F for p = .01

n\n, 1 Z, 5 4 5) 6 8 12 24 oo

1 4052 4999 5403 5625 5764 5859 5981 6106 6234 366
2 98.49 99.01 Mil YW. 99.3 a5 O95 Le. 99.4 9.50
3 34.12 30.81 29.46 28.71 28.24 PATON 27.49 27.05 26.60 Zone,
4 21.20 18.00 16.69 15.98 15.52 IhSy7 14.80 14.37 IE 3.46
5 16.26 S227 12.06 1S2 LOLO7, 10.67 10.27 9.89 9.47 9.02
6 Sele. 10.92 Dakss OnID OMe 8.47 8.10 TAZ 731 6.88
a5 ee OVS 8.45 7.85 7.46 7.19 6.84 647 6.07 5.65
_— _ iw) ron Go renWN 7.59 7.0 6.63 6.37 6.03 5.67 5.28 4.86
XI
©‘9 10.56 8.02 6.99 6.42 6.06 5.80 5.47 5.11 473 4.3]
(Continued)
332 << STATISTICS FOR THE SOCIAL SCIENCES

Table 10.1 (Continued)

n\n, 1 a 2in a a3 4 5 6 8 12 24 co
See: ee ee 2 ee Pe
10 1004, 756 655 S590 5.64 5.39 5.06 471 4.33 3.91
11 9.65 7.20 6.22 567 - 5:32 5.07 474 440 4.02 3.60
12 9.33 6.93 5.95 5.41 5.06 4.82 450 416 3.78 3.36
13 907 670 S74 520 486 462 430 3.96 3.59 3.16
14 8.86 6.51 556 5.03 469 4.46 414 380 3.43 3.00
15 8.68 636 542 489 456 432 4.00 367 3.29 2.87
16 8.53 6.23 520 ary "eae “420 ~ BO” GSS 3.18 2.75
17 840 6.11 ca ee A ce (I Ss 2.65
18 8.28 6.01 509 458 425 401 3.71 3.37 3.00 2.57
19 818 5.93 5.01 450 417 394 3.63 3.30 2.92 2.49
20 @10 585 404 (443) 490 “287 «356 320 2.86 2.42
21 8.02 5.78 487 437 404 3681 551 g07 280 2.36
22 794 5.72 4.82 431 3.99 3.76 3.45 3.12 2.75 2.31
23 788 566 476 426 394 3.71 3.41 3.07 27 2.26
24 7.82 5.61 472 4.22 3.90 3.67 3.36 3.03 2.66 Aa
25 Tit Sis "46S AIG. BBE HGS 3.32 2.99 2.62 2.17
26 772. S55 Ae 414 «= SRBE CSSD G2 2.96 2.58 2.13
27 768 549 460 411 3.78 3.56 3.26 2.93 2.55 2.10
28 JA 545 457 407 375 3:53 3.23 2.90 2.52 2.06
29 760 542 454 404 3.73 3.50 3.20 2.87 2.49 2.03
30 156 “539 481 402 370 347 ~ 3.17 2.84 2.47 2.01
40 731 518 431 $63 3.29 2.99 2.66 2.29 1.80
60 708 = ==408 413 $85 $34 312 2,82 2.50 2.12 1.60
120 685 479 3.95 348 317 2.96 2.66 2.34 1.95 1.38
20 6.64 460 3.78 332 3.02 2.80 Zot 218 1.79 1,00
Table 10.1 Continued—Critical Values of F for p = .001

n\n, 1 my 3 4 5) 6 8 12 24 oo

1 405284 500000 540379 562500 576405 585937 598144 610667 623497 636619
2 998.5 999.0 999.2 999.2 Oe) 9993 999.4 999.4 993° SINS:
S) 167.5 148.5 141.1 137 134.6 132.8 130.6 128.3 125.9 123.5
4 74.14 61.25 56.18 53.44 Silva 50.53 49.00 47.41 45.77 44.05
5 47.04 36.61 33.20 31.09 EMITS 28.84 27.64 26.42 25.14 23.78
6 Sys)! 27.00 23.70 21.90 20.81 20.03 19.03 A) GEO bas
7 Zaz 21.69 Si PAIS) 16.21 15.52 14.63 iol WEA fe} 11.69
8 25.42 18.49 15.83 14.39 13.49 12.86 12.04 HAUSIRS, 10.30 9.34
9 22.86 16,39 13.90 12.56 ill g/l 11.13 10.37 Doi 8.72 7.81
10 21.04 14.91 12755 11.28 10.48 D2 9.20 8.45 7.64 6.76
iH! 19.69 13.81 11.56 10.35 9.58 9.05 8.35 7.63 6.85 6.00
12 18.64 T2097, 10.80 9.63 8.89 8.38 poral 7.00 6.25 5.42
3) 17.81 i2eok 10,21 9.07 BIOS 7.86 Wevall 6.52 Syits' 4.97
14 17,14 11.78 OS 8.62 ee 7.43 6.80 6.13 5.41 4.60
15 16.59 11.34 9.34 8.25 Teo 7.09 6.47 5.81 Sle 4.31
16 16.12 10.97 9.00 7.94 Wee, 6.81 6.19 3) 4.85 4.06
17 IW 10.66 ‘Shiv 7.68 7.02 6.56 5.96 DIoZ 4.63 BRED)

(Continued)
One-Way Analysis of Variance p 333

n\n, i 2 3 4 5 6 8 12 24 oo
18 15.38 10.39 8.49 7.46 6.81 6.35 5.76 5.13 4.45 3.67
19 15.08 10.16 8.28 7.26 6.61 6.18 5.59 4.97 4.29 3.52
20 14.82 9195 8.10 FAG, 6.46 6.02 5.44 4.82 4.15 3.38
at 14.59 9.77 7.94 6.95 6.32 5.88 5.31 4.70 4.03 3.26
yo 14.38 9.61 780 681 6.19 5.76 5.19 4.58 3.92 3.15
23 14.19 9.47 7.67 6.69 6.08 5.65 5.09 4.48 3.82 3.05
24 14.03 9.34 7.55 6.59 5.98 5.55 4.99 4.39 3.74 2.97
25 13.88 9.22 7.45 6.49 5.88 5.46 4.91 4.31 3.66 2.89
26 13.74 9.12 V6 6.41 5.80 5.38 4,83 4.24 3.59 ZAO2
27 13.61 9.02 77, 6.33 5.73 5.31 4.76 4.17 3.52 2.75
28 13.50 8.93 7.19 6.25 5.66 5.24 4.69 411 3.46 2.70
29 13.39 8.85 7.12 6.19 5.59 5.18 4.64 4.05 3.41 2.64
30 13.29 8.77 7.05 6.12 5.53 5.12 4.58 4.00 3.36 2.59
40 12.61 8.25 6.60 5.70 5.13 4.73 4.21 3.64 3.01 2.23
60 EO y 7.76 6.17 Soll 4.76 4.37 3.87 Sol 2.69 1.90
120 11.38 7.31 5.79 4.95 4.42 4.04 3.55 3.02 2.40 1.56
ro) 10.83 6.91 a2 4.62 4.10 Saat 3.21 Dy WES 2A\3 1.00

SOURCE: Abridged from Table V of R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricultural
and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint of Pearson Education.
NOTE: Values of 7, and 7, represent the degrees of freedom associated with the larger and smaller estimates
of variance, respectively.

equal category sizes; this ideal is not necessary and not always possible.
Note, too, that though we need not find the category means for the sample
to calculate F using these formulas, we do so anyway to better understand
the problem we are working. Applying these steps to the first problem pre-
sented in this chapter,

x = Pro-Life Index
x, = Rural x, = Urban
160 100
130 110 Ay: by = Lb
150 130 H,: M,#
n,=4 140 n,=4 120 Noy = 4+4=8
> # = 580 N= 400

460
oo Ge le oe Os
Ny 4 n2 4

St = 80 + 200 = 1040
334 << STATISTICS FOR THE SOCIAL SCIENCES

25,600 10,000
16,900 12,100
22,500 16,900
19,600 14,400
Yox7= 84,600 D3 = 53,400

YXral = D7 + Dx? = 84,600 + 53,400 = 138,000

BOX 10.1

Procedure for Calculating One-Way Analysis of Variance

1. Calculate the total sum of squares (sum of squared deviations from


the grand mean), where

2
oe a (D0 Xtotal)
SS Total =a eg a
Total
2. Calculate the between-group sum of squares (sum of squared devi-
ations of category means from the grand mean), where

2
(Soar ft)?ef (3) Xeata)” Me
SS Between Sid eek Swe
MNeat.1 Mcat.2
p ?

. ohne au) (>-<Toral)


Mast cat. Total

Note: +--+ ++ means continue repeating the


()ox.,,)/2.4 for the subsequent categories, if any.
3. Calculate the within-group sum of squares, where

5S“Within SS
~ “otal DO geneen

4. Calculate the between and within mean squares, where

M SS Between oS Between
SBetween ee :
Af petween 0. of categories — 1

SSwithin SS Within:
Y
MSwithin =
Af Within (Mota) — 1)(no. of categories — 1)
One-Way Analysis of Variance » 335

5. Calculate the F ratio, where

jie MS Between

MSwithin
6. Use the table of F values in your textbook to test the F for significance.

Note: Numerator df = dfy..yeen = NO. Of categories — 1


Denominator df= Af icin = Mprai— 1) — (no. of categories — 1)

7. If F is significant and the number of categories is greater than two,


perform a post hoc procedure such as Scheffé’s test.
8. If F is significant and this is nonexperimental research, measure
association with the correlation ratio or 7, (These will be presented
in Chapter 13.)

Following the steps in Box 10.1:

1. We find the total sum of squares.

2 z
5) (S2 xtotal) (1040)
SS
— ) “Total
— ———Notal
= 138,000
5
— 8

1,600
= 158;000— =— = 138,000 — 135,200 = 2800

2. We calculate the between-group sum of squares. Note that the last


expression in this formula, (0X? pia)°/rorat» WAS already calculated in
step 1)

Ox)? . Of)? ~ Olea)?


Ss.
1) n2 Total

2 é 6,400 211,600
= o- fe = 135.200: 2 a ae y 135,200

— 84,100 + 52,900 — 135,200 = 137,000 — 135,200 = 1800


336 @ STATISTICS FOR THE SOCIAL SCIENCES

3. We find the within-group sum of squares.

SS, = SS, — SS, = 2800 — 1800 = 1000


4, We calculate the mean squares.

1800 1800
MG, ee teal = = = 1800
G@fp. 0.01 categories = 1. Z—1 1

SSw 1000
MSy = =
afy; (tora — 1) — (no. of categories — 1)

sy(8—1)—-(2—-1)
cl 1 Ut 2 gE1000
7-1
1000
6
ria
5. We calculate the F ratio.

MSz 1800
Pe = Us Oe a
~ MSw 166.67

6. We find the degrees of freedom.

af, = no. of categories -1=2-1=1

Afy = (Myra — 1) — (no. of categories — 1)


=(8-1)-(@-1)=7-1=6
We then find F..,,,.q at the .05 level from the first table in Table 10.1, going
along the top row to df, (1) and dropping down that column until it inter-
sects the row for df, (6):

| : < df
df. 1 161.40 |199.501
2 18.51
6 5.99

F critical, .05 level, df= 1 and 6.

SINCE Foainea 1S 10.80 and greater than 5.99, we reject i, Using: the
second table of Table 10.1, we repeat the procedure to find F ag ae tne
One-Way Analysis of Variance » 337

-01 level, which is 13.74 and greater than 10.80. Sincep is less than .05 but
greater than .01, we state our probability of falsely rejecting a true null
hypothesis asp < .05.
Before moving on, let us return to Table 10.1 for a moment. From time
to time, we will calculate a degrees-of-freedom figure that does not appear
in the table, such as df, = 7 or dfy, = 31. In such a case, we use the adjacent
row or column that has the higher value of F Thus, for df, = 7, we would
see which value was higher, F at df= 6 or F at df= 8. For example, if df, =
7 and df, = 6, we would have a choice between F.,,,,..,, df 6 and 6 (4.28), or
F citica OF 8 and 6 (4.15). We use the larger of the two (4.28) as our critical
value. If df, = 1 and df, = 31, we have a choice between F.,,,,..,,, df 1 and 30
(4.17), OF Foitica Af 1 and 40 (4.08). Again, we select the larger of the two
(4.17) as our critical value. If we are very close to rejecting H, using this
procedure but do not quite make it, our best bet is to use a computer pro-
gram that reports the exact probability.
We go back once again to our ANOVA problem for which we have now
rejected H, with a probability of error < .05. We may wish to summarize our
findings in what is called an ANOVA source table.

ANOVA source table A table summarizing the results of the main steps in the
ANOVA procedure.

Source SS df MS F Je
Total 2800
Between 1800 1 1800 10.80 <n (5)
Within 1000 6 166.67

Note that we usually do not report df;,,,, (in this case, —-1=8-1=7,
or MS,,,, Total? which is 6’) since neither was necessary for finding F

COMPARING F WITH t

At this juncture, let us pause to compare F to the two-sample ¢ tests


presented previously. Although they are treated in most Statistics texts as
separate procedures, they are in fact mathematically related. In the case
of a two-category independent variable such as the problem just com-
pleted, it turns. out that if the two-sample ¢ test, assuming equal population
338 <4 STATISTICS FOR THE SOCIAL SCIENCES

variances, had been calculated on the same data, then F = Tinnour.


problem, calculating the ¢ value would yield 3.287, and squaring that yields
10.80, the same as our F
If we may do either ¢ or F when comparing two groups, which is
preferable? Generally, the two-sample ¢ test is, for two reasons. First, unlike
ANOVA, it is possible to make a directionality assumption in a ¢ test’s 1.
If we had advance reasons to believe rural residents were more pro-life than
urban ones, we could have used the directional critical values of t and
reduced our probability of error by one half.
Second, ANOVA as presented here assumes equal population
variances. If there is reason to assume unequal population variances, a
t test formula must be used. In fact, it is somewhat ironic that the assump-
tion of equal population variances plays such a major role in the f test
procedure since in most actual applications of ANOVA, the researchers
merely make the equal variance assumption and do F without testing the
assumption. The reason is that statisticians consider F to be robust, a
term in statistics meaning accurate even when underlying assumptions
(such as equal population variances) are violated. This is particularly true
when all category sizes are the same. Despite this observation about & if
evidence suggests very unequal population variances, this procedure
should be avoided.

Robust Accurate even when underlying assumptions (such as equal population


variances) are violated.

Finally, note that just as was the case with the ¢ test, ANOVA assumes
that the populations from which the categories are drawn are normally
distributed along the dependent variable. In our samples, if category
sizes are sufficiently large, we may relax the normality assumption. In this
respect, ANOVA is the same as the ¢ test.

ANALYSIS OF VARIANCE WITH EXPERIMENTAL DATA

Let us now turn our attention to the training of advertising consultants


problem. Recall that Hy: May = Haight = Usavurday aNd that H,: there exists
at least one inequality that negates H,. We had 10 trainees each in three
separate classes.
One-Way Analysis of Variance » 339

ay Night Saturday
x1 = Xx, = Ne

S 5 5
5 4 i
5 4. 5 Ngai =, +N, +N,
? 3 4 =10+10+10
; 2
2
4
D
= 30
4
4 2 3
4 @) 3
3 1 3

10s gant ye Ok ae Ors Fi, 10: 2


a mat Dae = 26 ee = 38

x,=4.1 X, = 2.6 X= 8.8

> wea >. ed + Dy ot yx, = 41 + 26 + 38 = 105

xT= cS x=
25 25 25
25 16 25
25 16 25
25 9 16
16 9 16
16 4 16
16 4 9
16 4 9
2 1 9
4 0 4

ye 7 Oe eres Yin=?154

Dot ge tO 17 88 154 419

Following the steps in Box 10.1:

eae a1) ON BOS

= 419 — 367.50 = 51.50


340 << STATISTICS FOR THE SOCIAL SCIENCES

tN Pi Oxy i Cas i Cag)? Ol ar)


SSB
al nN2 3 NY

_ 41? 2 | 26) 2 | G8) 2 367.50


10 10

Besa
1681
10 .
676
a0) v
es Mes
1444
10 :
gr
= 168.10 + 67.60 + 144.40 — 367.50

= 1050510) = 367.50 == 12360

3, SSy = SS, — SS,= 51.50 — 12.60 = 38.90

4. SSB SSp ne 12.60


= 6.30
df, no. of categories—1 3-1

Vie SS W SS Ww
Af ~ (Ntoral — 1) — (no. of categories — 1)

u 38.90 38.90 _ 38.90


= 1.44
COIS G= 22 er

Ds MS 6.30
Goze ees, ee 4.375
MSyw —«1.44

Consulting Table 10.1 forFcritica , at df= 2 and 27:

F ica 05 = 3.35< 4.375 reject H,

F critical? Ol=540s4 575° “p05

POST HOC TESTING


In rejecting the null hypothesis, we conclude H,: in the population there
exists at least one inequality that negates H,. Though we can conclude that
in the population, there is a difference between student satisfaction and the
time that the class is offered, we cannot be sure exactly where that differ-
ence lies since only one inequality negates H,. The possibilities here are as
follows:
One-Way Analysis of Variance » 341

My Uy # Ls
My # Uy = Us
My= My # My
My = Us # Uy

In the context of this specific problem, {, = Uy # Ll, is probably illogical


since X,, is closer to X, than to X,, although sampling error could conceiv-
ably have yielded these means from a population where Lp = Uy # Ly is really
true. Also, in the context of sampling, we might have a situation where we
cannot reject the null hypothesis either between fy, and LU, or between LL,
and [U,,, but we can reject it between fl, and LL).
Why not run a series of two-sample ¢ tests between each pair of means
and see which null hypotheses could be rejected? The answer is that since
the ¢ test was predicated on the assumption that a single comparison of
means was to be made, if we make more than one comparison, the proba-
bility of a Type I or alpha error increases with each comparison.
However, there are a number of tests, known as post hoc tests of mul-
tiple comparisons, that control for such inflated alpha levels and enable
us to narrow our conclusion regarding exactly where these population
inequalities are. One such test is Scheffé’s test (pronounced as in French:
shef-FAY). There are other more powerful control procedures, but we pre-
sent this one because of its flexibility and robustness. It can be applied even
when the groups being compared have different sizes (some tests assume
equal 7s), and it is less sensitive to departures from normality and any
assumptions of equal population variances than are some other tests.

Post hoc tests of multiple comparisons Tests that enable us to narrow our
conclusion to specifically where these population inequalities are to be found.

Scheffé’s test A test that finds the critical difference between any two sample
means that is necessary to reject the null hypothesis that their corresponding
population means are equal.

Scheffé’s test finds the critical difference between any two sample
means that is necessary to reject the null hypothesis that their correspond-
ing population means are equal. If U, # Hy, how big must the difference
be between X,, and x,? This difference, Scheffé’s critical value, may
be calculated between each pair of means, and the actual sample mean
differences are compared to the critical values. If |x,-x, | for any two
categories, 7 and j, exceeds Scheffé’s critical value, we may reject H, and
conclude LM, # L,.
342 << STATISTICS FOR THE SOCIAL SCIENCES

cee LO LA LL LL LILLIE TET

Scheffé’s critical value The value in this test needed to reject the null hypothesis.
ANN LN NEAL ENC NTO ON A ALTE EDT NC CAA,

We begin by presenting the ANOVA source table for the problem just
completed.

Source SS df MS fi Dp
Total SSO
Between 12.60 2 6.30 4.375 < .05
Within 38.90 Zi 1.44

For any two categories, 7 and/, the following formula generates Scheffe’s
critical value.

1 1
OX; — Xj critical = £ |r Fonant¥Sx) (— i ~)
nm ny

From the source table, we see that df, = 2 and MS,,=1.44. The F_,,,,.. -05 level
from Table 10.1 is F jiica 05. @f = 2 and 27 = 3.35. Finally, since all of our cat-
egory ms are equal to 10, 7, =”, = 10. Thus, one critical value will apply to all
three mean comparisons. Had our category sizes been unequal, we would
have had to calculate a separate critical value for each pair of sample means.
Plugging into our formula,

ha : ; 1 1
(X; — Xj critical => |r nFnant¥Se) (— i —)
ni Nn;

ere
= 2)G.05)
(2)Go0)( (as)“(Zt =)

= ay 1.0290 = aL. oae

In other words, the absolute value of any pair of sample mean differences
must equal or exceed 1.389 in order to reject H,. Examining our sample X s,
we see the following:

Scheffé’s
jek |x,-%, | = Critical Value Conclusion

Mp =Hy |4.1-2.6] =15 > 1.389 RejectH,


ia aloe Oe < 1.389 Cannot reject H,
aie 126=38 kate < 1.389 Cannot reject H,
One-Way Analysis of Variance j» 343

Thus, although our overall F was significant, we have traced that fact
to the single explanation of an inequality between the day and night class
population means. We cannot conclude that the population mean for the
Saturday group differs from either the day or the night classes.
Suppose, however, that |z,- x, | had been larger than 1.389, and we
could also have rejected H,. Our conclusion would be modified: The signif-
icant F resulted from the difference between the night class’s scores, on one
hand, and the combined day and Saturday scores, on the other. Since the
day and Saturday scores are not significantly different, we might conclude
that, since the Saturday classes also met during the daytime, it is the day ver-
sus night difference that counts, regardless of which day or days of the week
that the day class is held.
As noted earlier, there are many post hoc and a priori tests other than
Scheffé’s. These include Duncan’s multiple-range test, the Student-
Newman-Keuls’s multiple-range test, the least significant difference test,
Tukey’s honestly significant difference test, the Bonferroni procedure, and
others. Because of space considerations, only the calculation of Scheffe’s
test is presented here. Some of the other tests require extensive calculations
or additional tables of critical values. However, if you have access to a com-
puter, use of the Bonferroni test is generally preferred over Scheffe’s test
because it is easier to reject the null hypothesis for each pair of differences.
Be aware, though, that there are circumstances where Scheffe’s test or
Tukey’s test may be a better one to use.” There is considerable debate over
which test is most appropriate for specific research situations. Consult an
advanced research design text to learn more about them.

COMPUTER APPLICATIONS

SPSS

ANOVA printouts resemble the source tables you have seen in this chapter.
In general, SPSS’s subprograms ONEWAY and ANOVA use the same terminol-
ogy used in this chapter. Table 10.2 shows the SPSS data list for the problem
we just completed. VARO0001 is the type of class with day coded as 1, night
coded as 2, and Saturday coded as 3. VAROO002 is the satisfaction rating.
To run the one-way analysis of variance, click on the menu bar as follows:
Statistics
Compare Means
One-Way ANOVA

use the
In the Dialog box, find VAROOQO002 on the left, click on it, and
left, click
upper arrow button to place it in the dependent list. Now, on the
on VARO0001 and move it into the factor list.
344 << STATISTICS FOR THE SOCIAL SCIENCES

Table 10.2

VAROOOOL VAROOO02

1 1.00 5.00
Zi 1.00 5.00
) 1.00 5.00
4 1.00 5.00
5 1.00 4.00
6 1.00 4.00
Zu 1.00 4.00
8 1.00 4.00
9 1.00 3.00
10 1.00 2.00
iil 2.00 5.00
iW; 2.00 4.00
13 2.00 4.00
14 2.00 3.00
15 2.00 3.00
16 2.00 2.00
17 2.00 2.00
18 2.00 2.00
i) 2.00 1.00
20 2.00 00
21 3.00 5.00
22, 3.00 5.00
25 3.00 5.00
24 3.00 4.00
z5 3.00 4.00
26 3.00 4.00
27 3.00 3.00
28 3.00 3.00
29 3.00 3.00
30 3.00 2.00

Now click on the post hoc button. From the list of tests, we will select
the Scheffé test. Click on this and then click continue. Now click the options
button and click on descriptive. We do not have to run the descriptive
statistics for this problem, but it is generally a useful option to run. Click
continue and when the Dialog box reappears, click ok.
In Table 10.3, the output for this run is reproduced. Note that first are
the descriptive statistics that we had opted to include. This is followed by
the ANOVA source table.
The Scheffé results are found in Table 10.4. Significant mean differences
are highlighted with an asterisk. Below that, two homogeneous subsets are
identified. Subset 1 contains the means for Groups 2 and 3, indicating no
One-Way Analysis of Variance p» 345

Table 10.3. Oneway

Descriptives
VAROOO02

95% Confidence Interval for Mean

Std. Std. Lower —_ Upper


N Mean Deviation Error Bound Bound Minimum Maximum

1.00 10 4.1000 9944 3145 3.38860 4.8114 2.00 5.00


2.00 10 2.6000 15055 4761 E5250 3.6770 OO 5.00
3.00 10 3.8000 1.0328 3266 3.0612 4.5388 2.00 5.00
Total 30 3.5000 1.3326 2433 3.0024 3.9976 .00 5.00

ANOVA
VAROO002

Sum of Squares af Mean Squares F Sig.

Between Groups 12.600 2 6.300 4,373 .023


Within Groups 38.900 Pai 1.441
Total 51.500 29

population differences in satisfaction when night students are compared to


Saturday students. Subset 2 indicates no population differences between
day and Saturday students. Note the absence of a subset containing the
means of Group 1 and Group 2. This parallels our earlier finding that there
is a significant difference between day and Saturday attendees.

SAS

Click as before:

Solutions
Analysis
Analyst

Enter the data just as in Table 10.2, with A being the category (either 1,
2, or 3) and B being the satisfaction rating. Then click

Statistics
ANOVA
One-Way ANOVA

(Shortcut: The second icon from the right at the top of the page,p = 05,
also clicks you into one-way ANOVA.)
346 << STATISTICS FOR THE SOCIAL SCIENCES

Table 10.4 Post Hoc Tests

Multiple Comparisons

Dependent Variable: VAROOO02


Scheffe

95% Confidence Interval

Mean Std. Lower Upper


(1) VAROOOOL (2) VAROOOO1 Difference (I-]) Error Sig. Bound Bound

1.00 2.00 1.5000* Oy 032 .1097 2.8903


3.00 .3000 Soe .856 —1.0903 1.6903
2.00 1.00 —=1,5000* Dar .032 —2.8903 —.1097
3.00 —1.2000 OT 101 = 2.9905 .1903
3.00 1.00 —.3000 ot .856 —1.6903 1.0903
2.00 1.2000 DoW, SOM = ASI0S 2.5903

Homogeneous Subsets

VAROOO02
Scheffe*

Subset for alpha =.05

VAROO001 N 1 2
2.0 10 2.6000
3.00 10 3.8000 3.8000
1.00 10 4.100
Sig. 101 856
NOTE: Means for groups in homogeneous subsets are displayed.

a. Uses Harmonic Mean Sample Size = 10,000.

*The mean difference is significant at the .0S level.

In the Dialog box, move B to the Dependent Variable box and A to the
Independent Variable box. At the bottom of the Dialog box is a button
labeled means. A new Dialog box will open.
It is here that you will select the post hoc test or tests you wish run.
Click on A in the box labeled Main Effects (on the left-hand side of the
screen). A menu of post hoc tests will appear. You may pick whatever test
you want. In this example, you would want to highlight Scheffé’s Multiple
Comparison Method. Then click on the add button just above the
effects/methods box, and the Scheffé test will be listed in that box. If Scheffé
is all you want, click the ok button.
One-Way Analysis of Variance p 347

To add additional tests, do not click ok. Instead, click on the down arrow
icon to the right of the box under Comparison Method, to be found near the
top of the screen. The menu of tests will appear. Click on the test to be
added, and the menu will disappear with the test you just selected listed in
the box below Comparison Method. Again click on A on the Main Effects
box and then click on the Add button. The new test is now found in the
effects/methods box under the Scheffé test. To add more tests, go back to the
down arrow icon and click it. Repeat the procedure by highlighting the third
test you want done and follow the same procedure as above.
When done, click ok to go back to the first Dialog box, and click
ok again to run the ANOVA. The output will appear on the screen as in
Tables 10.5 (ANOVA) and 10.6 (Scheffé),

Table 10.5 — SAS Analysis of Variance Output

The ANOVA Procedure


Class Level Information

Class Levels Values

A c L235

Number of Observations Read 30


Number of Observations Used 30

The ANOVA Procedure


Dependent Variable: B

Sum of Mean
Source DF Squares Square F Value 2p soi

Model 2 12.60000000 6.30000000 4.37 0.0226


Error 2) 38.90000000 1.44074074
Corrected 29 51.50000000
Total

R-Square Coeff Var Root MSE B Mean

0.244660 34.29453 1.200309 3.500000

Source DF Anova SS Mean Square F Value Pr>F

A 2 12.60000000 6.30000000 AST 0.0226


348 << STATISTICS FOR THE SOCIAL SCIENCES

Table 10.6 SAS Post Hoc Test (Scheffé’s Test)

Output: The ANOVA Procedure, Scheffé’s Test for B

The ANOVA Procedure

Scheffé’s Test for B

NOTE: This test controls the Type I experiment-wise error rate.

Alpha 0.05

Error Degrees of Freedom ra


Error Mean Square 1.440741
Critical Value of F 3.35413
Minimum Significant Difference 1.3903

Scheffe Grouping Mean N A

A 4.1000 10 1
A 3.8000 10 3
A
B
B
B 2.6000 10 2D

Excel

As with the ¢ test in Excel, the data are entered in columns, with column
A being the day class, column B being the night class, and column C being
the Saturday class (see Figure 10.1). Then click

Tools
Data Analysis
ANOVA: Single Factor

Highlight the input range as before. In this case, $A$1:$C$10 should


appear in the box, and make sure the Grouped by Columns button is
indicated. Click ok. The results also appear in Figure 10.1. No post hoc tests
are available with this routine.

TWO-WAY ANALYSIS OF VARIANCE


We have only scratched the surface of ANOVA, covering topics most ger-
mane to social scientists who will generally use nonexperimental techniques.
One-Way Analysis of Variance j» 349

Figure 10.1 Excel Data and Printout for ANOVA

Input:

5 5 5
5 4 5
3) 4 5
5 3 4
4 3) 4
4 Z 4
4 4 3
4 2 a
2) 1 3
2 0 Z

Output:
ANOVA: Single Factor

SUMMARY

Average Variance

0.988889

ANOVA

Source of
Variance P-value

Between Groups 4.372751 0.022642 | 3.354131

Within Groups 1.440741

In behavioral science applications, where experiments are more common, a


wide variety of elaborate advanced analysis of variance techniques are used
by researchers.
In two-way analysis of variance, we extend our model to include
a second independent variable. Suppose in our training example, we also
wanted to see if subjects’ satisfaction with the course could be explained
by the nature of their specialization within the firm. Suppose half of the
350 << STATISTICS FOR THE SOCIAL SCIENCES

subjects were employees whose missions emphasized advertising strategies


and issues, whereas the other half of the subjects were primarily media con-
sultants who were less issue oriented and more concerned with communi-
cations skills and use of mass media. Do the two groups respond similarly
to the scheduling of the course?

Two-way analysis of variance Analysis of variance that includes a second


independent variable.

Suppose the following pattern emerges:

Time of Class

Day Night Saturday

x= x %
> 5 2
5 4 5
for Advertising Strategists x = S 4 5
5 a) 4
4 a) 4
K=4.8 x = 3.8 6 = 416
4 2 +
4 2 3
for Media Specialists x = 1 2 3
5) 1 A
2 0 a
X = 3.4 xXx=1.4 x = 3.0

Assuming that F is significant when comparing our six means, we see


the same trend as before: Day and Saturday classes are preferred to
night classes for both advertising strategists and media specialists.
However, we note the consistently higher satisfaction of the strategists.
Their satisfaction is higher than their media colleagues, regardless
of the time of the class. We would need further research to determine
why these differences exist; perhaps media people need a different kind
of training program.
On the other hand, suppose the following pattern of scores had
emerged.
One-Way Analysis of Variance p» 351

Time of Class

Day Night Saturday


S2 I 2 Il be I

5
4
for Advertising Strategists x = 4
2)
2
Fe
are nee
Rem
Le
3| l| Uo ON <3 l| S ON ba I| Us OV

for Media Specialists x =

Renae ae
es
lee t:
Panties.
3 II aN OV | II 7 om | II aN Ss

Here, the strategists are unaffected by class scheduling. The sample


means are the same, 3.6, for each class. We could assume no difference in
the populations for the advertising strategists.
For the media specialists, assuming F is significant; not only are there
differences in satisfaction from class to class, but the day session is most
popular. Also, both day and Saturday mean scores for the media specialists
are greater than those for advertising strategists. Only the night class is
unpopular among the media people. Such a situation suggests many
possible follow-up studies.
Also, there are situations where the relationship between the depen-
dent variable and one of the independent variables is a function of the levels
of the other independent variable. These are known as interaction effects,
and two-way ANOVA also measures such effects.

Interaction effects Situations where the relationship between the dependent


variable and one of the independent variables is a function of the levels of the other
independent variable.

Social scientists have recently been doing greater numbers of experi-


ments than in the past. For example, one now finds structured simulations
352 << STATISTICS FOR THE SOCIAL SCIENCES

of decision-making processes under conditions varied by the experimenter


for different groups. Accordingly, two-way ANOVA techniques may soon
become as important in social science as they are in so many other fields.

CONCLUSION
We have now seen analysis of variance used in both experimental and
nonexperimental contexts and discussed Scheffé’s test as well as several
other procedures related to ANOVA. In Chapter 14, we will demonstrate
another context in which this procedure is applied, namely, as part of the
regression procedure. At that point, we shall have completed the process of
weaving together the two statistical strands—descriptive and inferential—
that have run through this text.

Chapter 10: Summary of Major Formulas

ipa... oi. =< Ce a ee ee ee


Analysis of Variance
2
s ‘ 2 (Total)
S Total — ie XToral cs
NTotal

Op eae rt (D> Xcat.2)° pe


SSBetween ae
NCat.1 NCat.2

a (oe See ex Xiu)”

| Nast Cat. N Total

SSwithin = SS Total = SSBetween

MSp ee SSBetween SSBetween


etween — “7 = P
Gf seween) @i@HOl Categones 1

MSyinin
.
= ‘———
SSWithin
dfwithin
= 99SSWwithi
Within
(ota — 1) — (no. of categories — 1)

MS Between
[Ml
MSwithin

= no. of categories — 1

= (Ntotal — 1) — (no. of categories — 1)


One-Way Analysis of Variance 353

Scheffé’s Test

For any two categories, 7 and 7, the following formula generates


Scheffé’s critical value:
ai | 1 1
i — Xj critical == (Gf g)(Feritical)(MSw) (~
l H =|
J

EXERCISES:
Exercise 10.1
Here is one of the example problems from Chapter 9. Perform ANOVA. Explain
why your conclusion differs from the one reached with the two-sample ftest.

Saw Commercial Did Not See Commercial

AO ee X=

Oo
Co7
C1
SON

©
Oo
Hm
“IAD

Exercise 10.2
_Ascale measuring support for increased gun control legislation (0 = no support to
5 = most support) is administered to random samples of urban, suburban, and rural
voters. Do the three population means differ in terms of support? If so, do Scheffé’s
test. What do you conclude? :

Urban Suburban Rural

xX, = xX, = X=

oawest
wk med
mek
CB
WT
&
Aw
Vik
354 4 STATISTICS FOR THE SOCIAL SCIENCES

Exercise 10.3
The same attitude scale used in the previous exercise is applied to random samples
of urban police officers, white-collar workers, and blue-collar workers. Do ANOVA,
and if the null hypothesis can be rejected, do Scheffé’s test. What do you conclude?

Police White Collar Blue Collar

R= xX, = x, =

5 4 1
4 3 3
6) 4 2
5 4 0
3 1 1
4 5
5

Exercise 10.4
For a random sample of Democrats in the U.S. House of Representatives, liberalism
scores were compared by region of the country. Find F and, if statistically significant,
do Scheffé’s test. What do you conclude?

Northwest South Midwest Far West

Sa X= x = xX, =
95 80 15 75
90 60 85 95
90 45 55 85
95 65 80 46
80 75 70 85

Exercise 10.5
For a random sample of 25 physicians, scores measuring support for a national
health care insurance program were compared by medical specialization.
Complete the resulting source table. What are your conclusions?

Source SS df MS E p
Total 32106.00 24
Between 21462.25 1 — — —
Within 10643.75 23 =

Exercise 10.6
From an experiment measuring the cognitive learning of students with learning
disabilities by various teaching strategies, ANOVA was run. Complete the source
table and state your conclusions.
One-Way Analysis of Variance j 355

Source SS df MS
Total 7339.84 —
Between 388.09 | —
Within — 23 =
~
Exercise 10.7
The same study as in Exercise 10.6 was done with students without learning
disabilities. Complete and interpret the source table.

Source SS df MS
Total _— 24
Between 4093.87 1 =
Within 2141.49 a a

Exercise 10.8
An ANOVA was run comparing political rights by GDP/capita (High, Medium,
Low, Very Low). What conclusions do you reach?

Sum of
Squares df Mean Square F Sig.
Between Groups 16.467 3 5.489 G7 t0n O01
Within Groups 10.083 16 .630
Total 26.590 19

Exercise 10.9
In a certain study, scores earned on a graduate school admissions test were
compared between those who had no formal preparation and those taking courses
designed to prepare students for the exam. Interpret the printout with regard to
statistical significance.

DEPENDENT VARIABLE: SCORE

SOURCE DP SUM OFSQUAKES MEAN SQUARE FVALUE. PRoft

MODEL 1 6.12500000 6.12500000 10.28 .0032


ERROR. 30 17.87500000 0,59583333
CORRECTED .31 24.00000000
TOTAL

Exercise 10.10
Here is a study of pilot reaction times under two different instrument panel configu-
rations. Interpret the printout.
356 << STATISTICS FOR THE SOCIAL SCIENCES

SOURCE DF SUM OF SQUARES MEAN SQUARE FVALUE PR>f

MODEL 1 3.380000000 3.380000000 7.14 0.0120


ERROR 30 14.19500000 0.47316667
CORRECTED 3! 17.57500000
TOTAL

Exercise 10.11
Using the computer, enter and run the following ANOVA data:

A. The pro-life scale by urban vs. rural. Compare your results to those presented
earlier in this chapter (pp. 333-337). Are they the same?
B.. the satisfaction score by type of class data (pp. 339-340 and p. 342). Compare
the ANOVA results. Also, in addition to running Scheffé’s test, run Bonferroni's
and a few other available post hoc tests. Do any of the results differ?
C. The data presented in Exercise 10.2. Compare the results to those you calculated
earlier.
D. The data presented in Exercise 10.3. Also compare to your earlier findings.
aati

NOTES
1. If you have read the previous chapter, you are already familiar with the use of the
F table. However, note that here we have tables for the .01 and .001 levels as well as for
the .05 level. In the last chapter, we used only the .05 level table.
2. J. Neter, W. Wasserman, and M. Kutner, Applied Linear Statistical Models:
kegression, Analysis of Variance, and Experimental Designs (Homewood, IL: Irwin,
1985), p. 584.
a

a ae :

ae i
CwAartee

moe Association
Agency Tables

eiee ee |
SEE ee

2 Dal & (etm TW Cine ps haat


0) SF) Aa US is .00 |: eteoe |S

a pf PRIA Oath. Se pirsra We


PMR AE 1009 Siva as 1 Ore Prd), eT bh
Seger i . » By
1) Pel hee (Pre ikea
(0 it ABT Vis iiteocwi
sor*s) Ww a4
ini be 7 vall*

eo Oars ww | has | 69)


aS Gala =
° 5
WV KEY CONCEPTS ¥
SY IEE BEN GER CY EEN GEL IIE RIE LES LIE EAL ESSELTE EE TEER CDNEELENLEL ELLELOD IEEE ELEN EEEDD SEIDEL PLS STE

measures of symmetric versus Cramer’s V


association asymmetric measures Kendall’s tau-b
concordant pairs of association Kendall’s (or Stuart’s) tau-c
discordant pairs proportionate reduction in Somer’s d
Yule’s O errors (PRE) measures Goodman-Kruskal’s
phi coefficient lambda symmetric uncertainty coefficient
Goodman and curvilinearity versus Goodman and Kruskal’s tau
Kruskal’s gamma linearity in tables correlation coefficient
Goodman and Pearson’s C—the association (correlation)
Kruskal’s lambda contingency coefficient matrix
DAMES LLC LIEBE
LYELLEMD SESLSE ES AeA RO AAAS NTI LEENA
CHAPTER

Measuring Association
in Contingency Tables

VW PROLOGUE

In Chapter 1, at the very beginning of the book, we were concerned with the
idea of association. Is the presence of rain associated with the presence of
clouds, and conversely, is the absence of rain associated with the absence of
cloudy conditions? In Chapter 6, we continued the development of this idea
with more refined contingency tables.
Specifically, association means that being in a specific category of one
variable (rainy, in the variable presence of rainfall) is associated with being in
a specific category of the other variable (cloudy, in the variable sky condi-
tions). So far, we have a good idea of what constitutes a perfect relationship
(always, when we have rain, we have clouds, and always, when we have no
rain, we have no clouds). But, of course, that relationship is not really perfect
in that sometimes we have no rain, but the sky is nevertheless cloudy.
We also have an intuitive sense of what is meant by no relationship.
If the same proportion of rainy days have clouds as have no clouds and
if the same proportion of sunny days have clouds as have no clouds, then
rain and clouds occur independently of one another.
Between no relationship and a perfect relationship, however, we obvi-
ously have some relationship, but how much? In this chapter, we address
the problem of how much. We do so with what are termed measures of
association.

p® 359
360 < STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION

Generating a table of percentages is a common technique for hypothesis


testing, but the results in a table are not always clearly understood and
sometimes seem inconsistent. Moreover, except for the extremes of perfect
relationship and no relationship, it is not possible to see the exact amount
of relationship that exists between the variables. Measures of association,
which often accompany tables, seek to resolve this problem by showing the
amount and nature of the relationship presented.
Measures of association are index numbers that generally range in
magnitude from 0 (measuring no association) to 1 (meaning perfect asso-
ciation). Often, these numbers are signed, with a positive sign (+) denoting
a positive relationship between the variables and a negative sign (—) indi-
cating an inverse relationship. There are dozens of such measures, each
designed for specific table sizes and levels of measurement of the vari-
ables. Most of these measures are patterned after the Pearsonian correlation
coefficient (Pearson’s r), which we will encounter in Chapter 13.

Measures of association Index numbers that generally range in magnitude


from 0 (measuring no association) to 1 (usually meaning perfect association).

Most of the computer programs that generate tables (crosstabs) for


social science applications also generate a series of measures of association.
Each measure generated may or may not be valid for that crosstab, depend-
ing on features to be discussed shortly.
In this chapter, four of the most generally encountered measures will be
presented in detail. They will be sufficient for use in any data situation you
are likely to encounter.

MEASURES FOR TWO-BY-TWO TABLES

In a two-by-two table, each variable is a dichotomy, a two-category scale.


Although these scales may either be nominal or ordinal levels of measure-
ment, it is not necessary with dichotomies to differentiate nominal from
ordinal levels. This phenomenon was mentioned in Chapter 2, where we
pointed out that since it is impossible to destroy the sequencing of a
dichotomy’s categories, even a logically nominal dichotomy can be treated
as if it were an ordinal level of measurement.
Consider Table 11.1, a two-by-two table where the respondent’s income
(high versus low) is related to the respondent’s level of participation in
fund-raising activities for a charity (high versus low). Note that the table
entries are frequencies, not percentages. Every measure of association
Measuring Association in Contingency Tables » 361

presented here is calculated from the frequencies, even if the final accom-
panying table is in percentages. For ease in visualization, let us assume also
that the marginal totals are equal in magnitude.

Table 11.1

Income

Participation High Low Total

High 50 0 50
Low 0 50 50
Total 50 50 100

Whenever both variables are ordinal, as is the case in Table 11.1, the
categories are listed so that the upper left-hand cell of the table will contain
those cases scoring high on both variables, and the lower right-hand cell will
contain the cases scoring low on each variable. Then, to the extent that the
relationship between the two variables is positive, the clustering of cases will
be on the main diagonal (upper left to lower right). If the relationship is
inverse, the upper-right to the lower-left cells, the off diagonal, will show the
clustering. In Table 11.1, all 100 respondents or subjects cluster on the main
diagonal, indicating a positive relationship. Had that relationship been
inverse, the clustering would have been along the off diagonal (high partic-
ipation with low income; low participation with high income).
When one or both of the variables is measured at only the nominal level,
we set up the table so that if the hypothesis is verified, the cases will cluster
on the main diagonal. Consider the example used in Chapter 7. Suppose our
hypothesis is that one’s attitude on the issue of U.S. intervention in Latin
America to stop cocaine production and distribution is related to one’s atti-
tude on the issue of use of the death penalty on convicted high-level drug
dealers (“kingpins”) inside the United States. Supporters of intervention
also support the death penalty (see Table 11.2).

Table 11.2
Attitude on Intervention

For Against Total


Attitude on Death Penalty

15 5 20
For
> IS 20
Against
Total 20 20 40
362 < STATISTICS FOR THE SOCIAL SCIENCES

Though not a perfect relationship, 30 of the 40 subjects studied were


consistent with the hypothesis, 15 for both intervention and the death
penalty and 15 against both intervention and the death penalty. The remain-
ing 10 respondents were inconsistent with the hypothesis, with 5 against
intervention but for the death penalty and 5 for intervention but opposed
to the death penalty. The measures of association we will calculate for
Table 11.1 will have positive signs, denoting that clustering was along
the main diagonal. A clustering on the off diagonal would yield a negative
measure of association: Intervention attitude would be related to capital
punishment attitude, but 7of in the predicted direction.

Yule’s QO

Yule’s Q is an easy-to-calculate and easy-to-interpret measure of associa-


tion for any two-by-two table regardless of level of measurement. To use the
formula for Q, we define the cell entries in our table with alphabetical letters
as follows:

G=15, 0=5 ¢=s5 anda Sis

For dable di:

a=50,0=0,¢
30, anda =S0.
Measuring Association in Contingency Tables » 363

The formula for Q is

_ ad — bc
ole mee
Noting that ad means a times d and bc means b times c, Q for Table 11.2
would be

_ ad—be _ (15)(15) — (5)(5) _ 225-25 200


Q= ad+be = +0.80
(15)(15) + (5)(5) 225425 ~+«250
The positive sign indicates that the relationship is in the direction pre-
dicted by the hypothesis. Since Q ranges from 0 (no relationship) to 1.00
(a “perfect” relationship) in magnitude, we may assume that a Q of .80 is
a fairly strong relationship.
The value of Q has a technical interpretation that can be briefly
explained at this point. This interpretation is based on pairs of responses.
Assuming that in Table 11.2, our hypothesis is that that those favoring inter-
vention will also favor the death penalty and those against intervention will
also be against the death penalty, any respondent who is either “for” on both
responses or “against” on both responses will be consistent with the hypoth-
esis. Such consistent pairs of responses are also called concordant pairs.
By contrast, pairs of responses inconsistent with the hypothesis—for inter-
vention and against the death penalty or against intervention but for the
death penalty—are called discordant pairs. Yule’s Q is the proportionate
excess of concordant over discordant pairs of observations. Said another
way, QO is the net probability of randomly selecting a pair of responses from
the table that is consistent with the hypothesis. In this instance, we have a
.80 net probability (an 80% chance) of randomly selecting a respondent
whose responses are consistent with the hypothesis.

Concordant pairs Pairs of responses that are consistent with the hypothesis.

Discordant pairs Pairs of responses that are inconsistent with the hypothesis.

Yule’s Q The proportionate excess of concordant over discordant pairs of


observations.

For the data in Table 11.1,

_ ad —be _ (50)(50) — (0)(0) _ 2500-0 _ 2500


Sag 08)
~ ad+bce (50)(50) + (0)(0) 2500+0 2500
364 << STATISTICS FOR THE SOCIAL SCIENCES

There is a perfect positive relationship between participation and income.


Since bc = 0, there are no pairs of responses in Table 11.1 that are inconsis-
tent with the hypothesis. Thus, we have a 100% chance (a perfect chance)
of randomly selecting a respondent whose responses are consistent with the
hypothesis.
Using the same variables as in Table 11.1, an example of vo relationship
(see Table 11.3) would be as follows:

_ ad —be _ (25)(25) ~ (25)(25) _ 625-625 _ 0 _,


Q= ad+bc (25)(25) + (25)(25) 625+625 +1250

Since ad = bc, consistent pairs equal inconsistent pairs, the numerator is 0


and the quotient is 0.

Table 11.3

Income

Participation High Low Total

High 25 2 50
Low 25 25 50
Total 50 50 100

A perfect inverse relationship, all pairs discordant (see Table 11.4),


would look like this:

_ ad—be _ (0)(0)— (50)(60) _ 0-2500 —2500 |


~ ad+bc — (0)(0) + (50)(50) =0+2500 ~—-.2500 at

Remember that a negative numerator divided by a positive denominator


yields a negative quotient. ;

Table 11.4

Income

Participation High Low Total

High 0 50 50
Low 50 0 50
Total 50 50 100
Measuring Association in Contingency Tables j 365

One problem with Q is that the presence of a zero in any cell causes the
final quotient to have a value of either +1.00 or -1.00 (i.e., £1.00, read “plus
Or minus One” or “positive or negative one”). In Tables 11.1 and 11.4, the
relationships were perfect (all cases conformed to the hypothesis), and Q
was +1.00 and —1.00, respectively. Sometimes, however, there are exceptions
to the hypothesis, but one cell entry of 0 gives Q a magnitude of 1 even
though the relationship between the variables is not perfect (see Table 11.5).

Wee oe) (2G) OCS). 1125-0 1250 1125


= +1.00
O= ad the 2625)\45) -)G5) 9 ti2540 1125420 ~ 1125

Despite a Q of 1.00, the relationship is far from perfect. There were


35 high-income respondents with low participation, a contradiction to the
hypothesis. Another measure, the phi coefficient (@), does not share that
particular characteristic with O and is sometimes preferred for that reason,
despite a slightly more complicated computation.’

Table 11.5

Income

Participation High Low Total

High 25 0 25
Low 25) 45 80
Total 60 45 105

The Phi Coefficient

For the phi coefficient, we use the same lettering format as was used
with QO.

a le b @+ bd)

S ie d (Gara)

(Ganon (Od) (GaeaGrteG)

The formula for phi (pronounced to rhyme with the word bee rather
than the word buy) is

ad — bc
0 Wate taat+ob td)
366 @ STATISTICS FOR THE SOCIAL SCIENCES

For Table 11.5, where O = +1.00,

Pf ph ee _ (25)(45) = (0)(35)
°= /a@tbyerdyatoe+a) »/(25)B0)(0)45)
125 1125
—————— a = +0.4841
/5,400,000 2323.79

Rounding back two decimal places, ¢ = +0.48

Phi An alternative measure of association for a two-by-two table that is


sometimes preferable to Q.

Now notice for Table 11.1, where OQ was also +1.00 but the relationship
was a perfect one, @ reflects that fact.

2 aad. = 0c — (0)(50) — (0)(0)


Y
~ Satby(et+d(atab+d) JS0)(50)(50)(50)
2500-0 _ 2500 |
+1.00
~ ./6,250,000 2500 —

For the data in Table 11.2, where Q was +.80, phi is

7 ad — bc ee =) Ce) a) Le)
i V(a+b)(c+d)\(a+cy(b+d) “ / (20) (20) (20) (20)

225 — 25 200
= = — = +0.50
/160,000 400

Note that in comparing one table to another, we always compare O to QO


or ® to @; never do we compare Q in one table to @ in another. Also, while
@ will not be misleading the way OQ can sometimes be, all measures of asso-
Ciation possess certain quirks or defects. This is because in summarizing a
table with a single number, some information is naturally lost. All we can do
to minimize this problem is make sure that we do not select a measure that
is going to mislead our readers.
Phi is actually a variation on a measure known as Pearson’s 7, the corre-
lation coefficient, which we will discuss in Chapter 13.
Measuring Association in Contingency Tables » 367

MEASURES FOR 1-BY-n TABLES


Whenever one of our variables has more than two categories (an n-by-1
table), our selection of a measure of association is complicated by the fact
that we must first determine each variable’s level of measurement. The
appropriate measure of association may then be selected. We learn two
measures next, Goodman and Kruskal’s gamma (y) for ordinal-by-ordinal
tables and Goodman and Kruskal’s lambda (A) for nominal-by-nominal
tables.’

Goodman and Kruskal’s Gamma (Y)

Gamma is a measure designed for ordinal-by-ordinal tables. It may also


be used when one of the two variables is a nominal dichotomy (which we
may handle as if it were ordinal). Yule’s Q is actually a special case of gamma
for a two-by-two table, but we treat it here as a separate measure for the sake
of simplicity and treat gamma as a measure for tables larger than two-by-two.
The formula for gamma is

Go =i
ee

Pop Py

[Link] P,, are first determined for the table and then gamma is calculated.

Gamma_ A measure designed for ordinal-by-ordinal tables that may also be used
when one of the two variables is a nominal dichotomy.

As with QO, gamma evaluates the net proportion of concordant pairs


of observations—those consistent with the hypothesis—as compared with
all pairs of observations. The definitions of concordant and discordant pairs
in a larger than two-by-two table are somewhat more complex than in the
case of Yule’s QO. Concordant pairs, P,, are ordered the same on each vari-
able, cluster around the main diagonal, and indicate a positive relationship
in the table. Discordant pairs, P,, are ordered higher on one variable than
the other, cluster about the off diagonal, and suggest an inverse relation-
ship between the variables. If P, is greater than P,,, the numerator is a posi-
tive number and so is gamma. If P, = P,, concordant and discordant pairs
balance, the numerator becomes 0, and gamma becomes 0. If P, is less than
P,, the numerator—and gamma—will be negative, and the relationship in
the table will be inverse.
368 4 STATISTICS FOR THE SOCIAL SCIENCES

y=+1.00 vy=—.94

Instead of using formulas to find P, and P,, we use an algorithm, a set


of instructions that applies to tables of any size. In Table 11.6, we examine
fund-raising participation versus income, with each variable measured as a
three-category ordinal scale.

Table 11.6

Income

Participation High Medium Low

High 5 1 1
Medium 2 2
Low 0 1 4

To calculate P,, we begin by generating a series of subtables from


Table 11.6. To do this, we begin in the upper left-hand cell. We take note of
that cell entry and create a subtable containing everything that is both below
and to the right of the cell entry. We ignore any marginal totals in the table
(thus they are left off in Table 11.6). We can also skip any cell whose entry is
zero as well as any cell that has nothing below it and to the right. For the
upper left-hand cell, we get the following:
Measuring Association in Contingency Tables » 369

When we come to the low-income, high-participation category, we see that


there is nothing below it and to the right, so we ignore it.
Having exhausted the top row, we move to the far-left cell of the next
row down and begin again.

Moving one cell to the right:

Once again, when we move to the far-right cell on that row, we find nothing
below it and to the right, so we ignore it. This brings us to the far-left cell of
the bottom row. Noting that for the entire bottom row there is nothing
below and to the right, we have no more subtables to prepare. Thus, we
have generated the following:

Next, we add the numbers in each subtable and multiply the sum by the
original cell entry above and to the left of the subtable.

5(44+2+1+4)=5(1) =55
12+4=16) = 6
2(1 + 4) =2(5) =10
4(4)=4(4) =16

The sum of the four resulting products is P..

55
6
10
+16
P,=87
To find P,, we go “through the looking glass.” We start at the upper
right-hand cell of the table and move /eft on the row. This time we look to
see what is below us and also to the /eft. In short, we follow in reverse the
procedure used to find P,
370 4 STATISTICS FOR THE SOCIAL SCIENCES

4* Z

fe
(*We could have left this subtable out since there is only a zero below
and to the left of the cell entry.)
We add the numbers in each subtable, multiply the total by the cell
entry above and to the right of the subtable, and add these products to
find P,. These calculations are presented in the order that the subtables
were generated, moving right to left through the above subtables.

1264.2Ot =e 7
12+ 0) =12)= 2
20.2 2c
= 4(0) I lo
4(0)
P d =i)

Since P,= 87 and P, = 11,

_Ps—Pqg 87-11 _ 76
if = 10,7759 = +0./6
Peace, = Syed, 3G

Gamma is inappropriate for nominal-level data since it assumes that the


table is composed of two variables, each with categories appearing in logical
sequence, as was the case in Table 11.6. With nominal variables, there is no
logical sequence to the categories. Accordingly, for nominal-level data, we
need a measure of association that will remain the same for any given table
no matter how the categories are ordered. To illustrate the problem, let us
recast in Table 11.7 the data from Table 11.6 but with the income categories
no longer in logical sequence.

Table 11.7

Income

Participation High Low Medium

High 5 1 1
Medium 2
Low 0) 4 i
Measuring Association in Contingency Tables » 371

Let us calculate gamma for this table, For ie

5 2 2

5(2+4+4+1) =5(11) =55


Ee ENC es
2(4+1) =2(5) =10
2(1) =2(1) =+2
Pea
For PF:

4 1

[Lo
1(2+2+0+4)=1()= 8
1(2+0)=1(2)= 2
4(0 + 4)=4(4) = 16
2(0)=2(0)=+0
P,=26
Therefore,

P;—Pg 72-26 46
if = = = = +0.469 = +0.47
Pepig T2720 98 i i

Compare this result to the gamma of 0.78 obtained from Table 11.6.
A measure of association designed for nominal-by-nominal data would
have yielded the same result for either table; it would be insensitive to the
ordering of the categories.

Goodman and Kruskal’s Lambda (A)


Lambda is designed for a table where at least one variable is nominal
and is not a dichotomy. If gamma is inappropriate for the given table, we
may use lambda. Since lambda is designed for nominal-by-nominal tables, it
carries no sign (+ or —). Directionality of the relationship must be explained
by reference to the variables in the table.
372 < STATISTICS FOR THE SOCIAL SCIENCES

a eeccemmmmmmmmnnasmamnamneeeermenanaemmmmnnmememmmmnammmnemmmnnmmemmmmmnnmnntl

Lambda Designed for a table where at least one variable is nominal and is not
a dichotomy.

Unlike the other measures discussed in this chapter, lambda is an


asymmetric measure of association, meaning that the value of the
measure depends on whichever variable was dependent. In fact, there are
actually three lambdas for any given table: one for when the variable whose
categories are the table’s rows is the dependent variable, one for when the
variable whose categories are the table’s columns is the dependent variable,
and one an average of the first two. All our previously discussed measures
were symmetric; there was only one measure for any given table, and it
did not matter which variable was dependent or independent. In Table 11.6,
for instance, gamma was +.78 whether we were explaining income based on
participation or whether we were explaining participation based on income.
With lambda, we have to decide iz advance what the dependent variable is,
income or participation.

Asymmetric measure of association A measure whose value depends on


whichever variable was dependent.

Symmetric measure of association A measure whose value does not depend


on whether the variable was dependent or independent.

Lambda—Column Variable Dependent

In Table 11.8, we are studying a group of recent immigrants from the for-
mer U.S.S.R. who have come to North America. We have a three-category
nominal scale for occupation against a three-category nominal scale for
nationality grouping in the former Soviet Union. Gamma for this table is
inappropriate; therefore, we calculate lambda.

Table 11.8

Nationality

Occupation Russian Ukrainian Belorussian Total

Professional 10 1 20 45
Blue-collar 20 15 5 40
Farmer 10 20 5 35
Total 40 50 30 120
Measuring Association in Contingency Tables » 373

Let us first assume that the dependent variable is nationality, and we


wish to predict nationality based on knowledge of the respondent’s occu-
pation. To find the appropriate lambda, we first determine two figures,
called EZ, and E,. Lambda is based on a strategy for predicting respondents’
scores on the dependent variable. To find £,, we play a game in which we
know initially only the marginal totals of the dependent variable: We do not
know the data in the cells of the table itself,

Nationality
Russian Ukrainian Belorussian | Total
40 50 30 | 120

The strategy is to place all subjects in the largest category of nation-


ality and see how many assignment errors we will make. The largest
nationality category is Ukrainian, with 50 of the 120 respondents so listed.
If we were to put all 120 subjects in the Ukrainian category, we would
have correctly assigned the 50 respondents who actually were Ukrainian
but would have incorrectly assigned 70 people, the 40 Russians and the
30 Belorussians. The total number of assignment errors, 70, is E,. Note
that we make fewer errors assigning all 120 to Ukrainian than we would
make assigning them to either of the other two categories of the depen-
dent variable. If we assigned everyone to the Russian category, we would
incorrectly be assigning 50 Ukrainians and 30 Belorussians to the wrong
category, and we would make 80 assignment errors. If we put all 120 in
the Belorussian category, we would correctly assign 30 but incorrectly
assign 90 to the wrong category. Thus, the lambda strategy is to put every-
one in the /argest category and then tally up the number of error of
assignment made.

E, = 40 Russians + 30 Belorussians = 70 total errors

The logic behind lambda is that the greater the relationship between
the two variables in the table, the fewer will be our assignment errors
when we know the information in the complete table. £, is the total
number of assignment errors made when we know all of the data in the
table. In effect, we assign each respondent to a category of the dependent
variable based on knowledge of that respondent’s category of the inde-
pendent variable. We assign by each category of the independent variable,
keeping tabs of the errors in assignment made. When we are done with
all categories of the independent variable, we add up all the errors made
tO. Set E,.
374 @ STATISTICS FOR THE SOCIAL SCIENCES

Let us take the first category of the independent variable.

Nationality
Occupation Russian Ukrainian Belorussian | Total
Professional 10 15 20 | 45

The largest nationality category for the professional group is Belorussian.


If we assign all 45 professionals to Belorussian, we will correctly assign
20 people, but we will incorrectly assign 10 Russians and 15 Ukrainians.
Thus, we will make a total of 25 assignment errors for that category of our
independent variable. Now let us look at the next occupational category.

Nationality
Occupation Russian Ukrainian Belorussian | Total
Blue-collar 20 15 5 | 45

Here, the largest nationality category is Russian, with 20 of the 40 blue-


collar workers being Russians. If we put all 40 in the Russian category, we
will correctly assign 20 but incorrectly assign the 15 Ukrainian and 5
Belorussian blue-collar workers. Thus, 20 errors are made for this category
of occupation. For the last occupational category:

Nationality
Occupation Russian Ukrainian Belorussian | Total
Farmer 10 20 5 | 35

Here the largest category is Ukrainian. Putting all 35 farmers in the Ukrainian
slot, we will correctly assign 20 but incorrectly assign 15 (10 Russians and
5 Belorussians).
Now we tally up:

Occupational Category Assignment Errors Made


Professional 25
Blue-collar 20
Farmer 15

Total Eo

We made 70 prediction errors (Z,) knowing only the marginal totals


for nationality. Knowing all the information in the table, we reduced the
number of assignment errors to 60. Lambda is the proportionate reduction
Measuring Association in Contingency Tables we 375

in error (PRE). We find the total reduction in errors (70 — 60 = 10) and
express it as a proportion of £,. Thus,

Ey
— b> 7O— 60 ~ 10
Xr = = = — = .1428=
.14
By 70 70

Knowing the data in the table, we reduce our assignment errors by a


proportion of .14 or 14%.

Proportionate reduction in error (PRE) The reduction in assignment errors


when we know all the frequencies in the table, rather than just the totals,
expressed as a proportion of the errors made when knowing just the totals.

Lambda—Row Variable Dependent

To find the lambda where occupation is dependent, we follow the same


procedures but work column by column.

Occupation Total
Professional 45
Blue-collar 40
Farmer a5
Total 120

If we put all 120 respondents in the largest occupational category,


we will correctly assign 45 people but incorrectly assign a total of 75 people
(40 blue collar and 35 farmers). Thus, £, = 75.
We now go column by column through the table to find £;.

Occupation Russian Ukrainian Belorussian


Professional 10 il) 20
Blue-collar 20 LS 5
Farmer 10 20 5
Total 40 50 30
E, = Assignment Errors = 20 a 30 - 10 = 60

If we put all 40 Russians in their largest category of occupation


(blue collar), we will be right for 20 but wrong for another 20 (10 profes-
the
sionals and 10 farmers). Thus, we make 20 assignment errors here. For
Ukrainians, the largest category is farmer. Calling all 50 Ukrainians farmers,
376 << STATISTICS FOR THE SOCIAL SCIENCES

we will be right for 20 but wrong for 30. For the Belorussians, the largest
category is professional. Putting all 30 Belorussians in that category results
in 10 assignment errors. Thus,

E,=20+
30+ 10=60

Since Eb. was 7);

Av
_R=— 6-6
———.
15 _
= — = .20
By ie) ie.

Thus, we get a 20% reduction in error when predicting occupational type,


knowing respondents’ nationality categories.
Sometimes, we may have both lambdas but have no basis for calling one
or the other variable the dependent variable. In such instances, we can
“average” the two lambdas to get a lambda symmetric. Adding the two
lambdas together and dividing by two, we get

F 14+ .20 .34


symmetric — 50, a en

Note also that, in general, lambdas tend to produce lower numbers than
the other measures of association covered in this chapter. For the problem
in Table 11.6 where gamma was .78, we could calculate lambdas for the table
and compare them to the gamma. Although this is an ordinal-by-ordinal
table, finding lambda 7s permissible since it assumes a lower level of mea-
surement (nominal by nominal). However, it is not appropriate to calculate
gamma for a nominal-by-nominal table, since gamma assumes a higher level
of measurement. Let us find lambdas for Table 11.6.

Lambda symmetric The average of two lambdas.


LES BLISTER SELENE

Income

Participation High Medium Low | Total — Assignment Errors


High 5 1 fae Fs 2
Medium p) 4 2 | 8 4
Low 0 1 4 Ail aes 1
Total 7 6 5 | 20 Ley

For predicting income from participation, we assign all 20 respondents


to the largest income category. Both high income and low income have
Measuring Association in Contingency Tables 377

marginal totals of 7; it does mot matter which one we use. Either way,
we correctly assign 7 and incorrectly assign 13. Thus, £, = 13. Going row by
row, for the 7 respondents with high participation, the biggest category
of income is high. Putting all 7 in high income, we correctly assign 5 but
incorrectly assign 2. For the 8 medium-participation respondents, the
largest income category is medium. Assigning all 8 to medium income, we
make 4 errors. Finally, putting all 5 with low participation in the biggest
income category, low, we make one error. Thus,

_Fi-&, 13-7 6
r = .46
pt eS ae

When we treat participation as the dependent variable, we first assign all


20 to the largest marginal category for participation, medium. We correctly
assign 8 but incorrectly assign 12; E, = 12. Going column by column, we
note the following assignment errors:

Participation High Medium — Low | Total


High 5 1 1 | 7
Medium 2 4 2 | 8
Low 0 eS
Total yi 6 wi | 20
E, = Assignment Errors = z ae Z 7 ee 7

E,=24+2+3=7

my Nee = ey = 0
IZ IW

Therefore,

46 + .42 88
Asymmetric = roa = oy = .44

Recall that the same table produced a gamma of .78.

CURVILINEARITY

Sometimes, when gammas and lambdas are calculated for the same ordinal-
by-ordinal table, we get seemingly contradictory results. Consider Table 11.9.
We begin by calculating gamma.
378 @ STATISTICS FOR THE SOCIAL SCIENCES

Table 11.9

Income

Radicalism High Medium Low Total

High 15 0 15 30
Medium 5 5 5 15
Low 0 15 0 15
Total 20 20 20 60

Zany
D6 Ses + OPS Besa
5(15 + 0) = 5(15) = ne:
P, = 450

Forr,

ENE
DGecr s+ Ue 15) =1505) S375
5(0 + 15) = 5(€15) = 75
P, = 450
Therefore,

Ps—Pa 450-4500
y= = — =—_=0
Ps+Pa 450+450 900
Using gamma, we were unable to find a relationship in the table. But
note what happens when we find lambda. For income dependent,

E,
=40
E,=15+10+0=25
Fo =fy 40—=25 “15
A= —— = = SS 6
Le 40 40 ——
Measuring Association in Contingency Tables » 379

For radicalism dependent,

E, = 30

Fo=5+5+5=15

A
_Bi-B, 30-15 15 = 50)
ae ee 50
Therefore,

Asymmetric =

For lambdas, a .44 is quite large. This suggests that the two variables are
related, even though gamma was 0. This apparent contradiction has to do
with the mature of the relationship between the variables. Gamma is sensi-
tive only to Jinear relationships, where the cases in the table cluster along a
straight diagonal line, either the main diagonal or the off diagonal. The clus-
tering in Table 11.9 is not in a straight line but rather is a curve.

High Medium Low

High 15 0 15
Medium 5 5 5
Low 0 15 0)

Note in particular where the largest cell entries in each category fall.

High Medium Low

Medium 5 5 5
Low 0 0

We call such a relationship curvilinear to differentiate it from a straight-


line, or linear, relationship. In the case above, high levels of radicalism are
associated with both high and low income, and low levels of radicalism are
associated with medium income levels.

Curvilinear versus linear relationship A curved versus a straight-line relationship.


380 << STATISTICS FOR THE SOCIAL SCIENCES

Table 11.10 gives another example of curvilinearity. The curve is reversed,


with both high and low income associated with /ow traditionalism and mid-
dle income associated with high traditionalism. If we calculate gamma, we
will find it to be 0, whereas lambda is .50 whichever variable is dependent.

Table 11.10

Income

Traditionalism High Medium Low Total

High 0 5 0 5

Medium 0 a 0 5

Low 0 20
Total 10 10 10 30

If a relationship is curvilinear, a measure of association designed for


nominal-by-nominal tables is more likely to indicate that fact than an
ordinal-by-ordinal measure. This is because such a measure is sensitive to all
relationships, not just those clustering along a diagonal. Accordingly, if you
have an ordinal-by-ordinal table and you calculate a gamma that is small, it
is a good idea to inspect the table again. If curvilinearity appears to exist in
the table, do a nominal-by-nominal measure such as lambda.

OTHER MEASURES OF ASSOCIATION

In addition to the measures discussed above, there are dozens of others. In


the two most common sets of statistical computer programs for the social
sciences, SPSS and SAS, crosstab programs calculate phi, lambda (asymmet-
ric as well as symmetric), and gamma. (Remember that for a 2 x 2 table,
gamma is the same as Yule’s Q. Thus, all the measures discussed here are
generated by these two computer programs.) In addition, SPSS and SAS
generate other measures. For our purposes, it is only necessary to be aware
of the existence of these measures and their approximate equivalence to the
ones we have learned.
Pearson’s C (the contingency coefficient) and Cramer’s V are
similar to the phi coefficient but apply to tables larger than 2 X 2 as well. C
and V for a2 x 2 table are the same as phi. (It is possible, using a different
formula than the one we presented, to calculate phi for a table larger than
Measuring Association in Contingency Tables ®» 381

2 X 2, but Cand V are easier to interpret and, therefore, preferable to phi if


the table is larger than 2 x 2.)
RRS SOTTO ALISO LAY ESSE ELAS OSES EB OEE ON ETSI OMENS EE

Pearson’s C and Cramer’s V_ Measures that are similar to the phi coefficient but
are more accurate whenapplied to tables larger than 2 x 2.
ESTES TAE MLL BEETS AIO LTRS ETE LER OCREEAENL
IE AR ISOEES EE OER TES SOIT

Kendall’s tau-b and Kendall’s (or Stuart’s) tau-c are similar to gamma
and, like gamma, are symmetric measures of association. If the number
of rows differs from the number of columns, tau-c is preferred to tau-b. If the
number of rows equals the number of columns, tau-b is preferred.

Kendall’s tau-b an Kendall’s (or Stuart’s) tau-c Measures ao are tne to


gape: and arep yeameic measures of association.
esas SURES
EERO ESTEE ESE TER ESEEH
EERIE SIE SOLE TNS SEES

Somer’s d is similar to gamma, but unlike gamma, it is asymmetric, as


was the case with lambda. Goodman-Kruskal’s uncertainty coefficient
is functionally similar to Goodman and Kruskal’s lambda.
ss ER LACH RON

Somer’s d Measure that is similar to gamma but is asymmetric.


Goodman and Kruskal’s uncertainty coefficient Measure that is functionally similar
to Goodman and Kruskal’s lambda.
£8 TLESPE ESLER SALLE SESS SUITES SELLS SSIES LTE SLEEVE,

Finally, all of the above asymmetric measures may be made symmetric


by “averaging” the two asymmetric measures for a given table, the way we
did with lambda.
Figure 11.1 provides a flowchart for determining the appropriate measure
of association for various tables.
Note one additional measure in the figure: Goodman and Kruskal’s tau (t),
which is similar to lambda but uses a different assignment strategy.
L689 SSN EDOMLLISSSEREL NSS MME SEES SATE SELLA ES SSCL AIRES

Goodman and Kruskal’s tau. A measure similar to lambda.


MUSES
LEMURS SESE SEE AMEE SEDER SMES SABER ad

INTERPRETING AN ASSOCIATION MATRIX

When a common measure of association or, as we will see later, a correla-


tion coefficient is calculated for each pair of a series of variables, the results
are often presented in a special table known as an association matrix (or, if
applicable, a correlation matrix). Consider the example below where there
are six variables and gamma is calculated between each pair.
382 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 11.1 Decision Flowchart for Measures of Association

Is the table a 2x2 table?

Fa pms
Yes, 2x2 No, greater than 2 x2

Yule’s Q Are both variables ordinal?


9
(Note: For this purpose
consider a// dichotomies
to be ordinal level!)

at ee ee
Yes No, No,
Both One
Nominal Nominal
and One
Ordinal

Gamma (y)
Kendall’s tau (t)
Somer’s d

Goodman-kKruskal’s
lambda (A)
Goodman-kruskal’s
tau (T)

Goodman-kKruskal’s
Uncertainty Coefficient

©, C, and V

The following matrix (these are gammas) resulted from a study of the
costs of vandalism in student dormitories in a public university. The variables
are as follows:

Damage: Dollars of damage in a dorm per student


GPA; Mean grade point average of all dorm residents
% Male: Percentage of dorm residents who are male
% in State: Percentage of dorm residents whose families reside in state
Age: Mean age of dormitory residents
Beer: Mean number of cans of beer consumed by the residents in
one school year
Measuring Association in Contingency Tables » 383

Data were collected on 20 dormitory units, with each variable recorded


as high, medium, or low and the dormitory assigned to the appropriate
ranked category.

Damage =
High
Medium 10
Low )
Total 20

Thus, out of 20 dormitories studied, 5 had high levels of damage, 10 had


medium levels, and 5 had low levels. These three-category grouped ordinal
variables were then used to form 3 x 3 tables (one for each pair of variables),
and from each table, a gamma was calculated. The results are presented in a
form resembling a distance table on a road map. To find the gamma between
any pair of variables, find one variable on a row and the other on a column and
look to see what coefficient appears at the intersection of the row and column.
Notice two aspects of Table 11.11. First, the “main diagonal” contains only
1.00 where each variable’s row intersects the same variable’s column. The
gamma (or any measure) between a variable and itself is always 1.00, a perfect
association. A second aspect is that everything above and to the right of the
main diagonal is a mirror image of what is below and to the left of that diago-
nal. To find the association between damage and GPA, we can either look at the
intersection of the damage row and the GPA column (finding a —.74) or look at
the intersection of the damage column and the GPA row (also finding a —.74).

Damage GPA
~ Damage 1.00 —.74
GPA —.74 1.00

Table 11.11

Damage GPA % Male % in State Age Beer

Damage 1.00 —.74 .68 O1 —.66 85


GPA —.74 1.00 —.66 .02 50 =06
% Male .68 —.66 1.00 =05 01 50
% in State O01 02 05) 1.00 02 .20
Age —.66 SO =O .02 1.00 O1
Beer 85 —.66 50 .20 01 1.00

The two intersections have the same number in them since gamma is
symmetric. Had we used an asymmetric measure such as lambda and desig-
nated the variables in the rows as independent variables and the variables
384 < STATISTICS FOR THE SOCIAL SCIENCES

in the columns as dependent variables, the number at the intersection


of the damage row and GPA column would be the lambda for predicting GPA
from damage. The lambda for predicting damage from GPA would be at the
intersection of the GPA row and damage column. The two lambdas would
not necessarily be the same, as was the case with the two gammas.
When we have as our coefficients gammas or other symmetric
measures, we often save space by eliminating the coefficients above and to
the right of the main diagonal since they are redundant (see Table 11.12).
If you look for a gamma and find only a blank space, such as the Damage
row and the GPA column, reverse the row and column. Going down the
Damage column to the GPA row, we find a gamma of -.74.

Table 11.12

Damage GPA % Male % in State Age Beer

Damage 1.00
GPA S74 1.00
% Male .68 —.66 1.00
% in State O1 02 =.,039 1.00
Age —.66 50 ==) .02 1.00
Beer 85 —.66 50 .20 O01 1.00

Now we can select some variable as a dependent variable and see which
of the five remaining variables are associated with it and are, therefore, plau-
sible independent variables.
What are the likely characteristics of the students in a dorm in which
high damage was recorded? Damage is the dependent variable. We look at
Table 11.12 to see the gammas between damage and the other variables. In
this case, they are easily found in the Damage column. Looking down that
column, we see the following: Damage 1.00 (obviously), GPA —.74, % Male
.68, % in State .01, Age —.66, and Beer .85. It is often useful to list these
beginning with the highest positive value, working down toward zero, and
then out again to the largest negative coefcient:

Damage versus Beer 85


Damage versus % Male .68
Damage versus % in State .O1
Damage versus Age —.66
Damage versus GPA —.74

Beer consumption has the highest positive gamma with Damage (.85).
The greater the beer consumption, the greater the damage. The .68 with %
Males suggests that the greater the percentage of males in the dorm, the
Measuring Association in Contingency Tables » 385

greater the damage. The gamma of .01 between % in State and Damage
suggests little association, so we exclude that variable. Age and GPA are
inversely related to Damage. The greater the damage, the younger the
students and the lower their GPA.
To summarize, dorms with high damage rates tend to be those with
higher levels of drinking, greater proportions of male residents, younger
students, and students with low GPAs.
What would be the “ideal” composition of residents if one wanted
to minimize damage? To answer this question, simply reverse the conclu-
sion to the first question. Dorms with low levels of damage would have low
rates of beer consumption, a higher percentage of female residents, older
students, and students with higher GPAs.
What are the characteristics of a dormitory whose residents have rela-
tively high GPAs? Try to answer this one on your own.

Computer applications will be addressed in the following chapter.

CONCLUSION
As mentioned earlier, since association measures Summarize entire
crosstabs, the nuances found in the tables often are not reflected in the
association measures generated from them. Measures of association are
also dependent on the size of the table, the levels of measurement
of the variables, and whether or not the relationship in the table is linear.
Nevertheless, though they must be used with care, these measures
of association are valuable tools. The information in Table 11.12, for
instance, summarizes 15 meaningful cross-tabulations and can be gener-
ated easily by a computer. For this reason, association matrices are and
will continue to be valuable tools for data analysis.
We return to the topic of association in Chapter 13, which deals with
correlation-regression analysis.

Chapter 11: Summary of Major Formulas

Two-by-Two Tables

a | b | (G0)
Cc d (C+a)
(ao) | O+¢@) (at
ice Co)

(Continued)
386 @ STATISTICS FOR THE SOCIAL SCIENCES

(Continued)

ad — bc ad
— bc
cup etam re J(atb)(c+d)(a+c(b+d)

n-by-n Tables

P= Py Ei Es
Y= = Sa
P.-+-P, Ey

Consult the body of this chapter for the calculation of the intermediate
values needed to find gamma and lambda.

EXERCISES
Note: These exercises refer to the tables in the exercises for Chapter 6, not to the
tables in the body of that chapter.

Exercise 11.1
Calculate Q and ¢ for the following table:

Exercise 11.2
The same table is reproduced below, taken from an SPSS run of data from a
sample of 19 countries, not the same as those in the sample used in Chapter 6.
VARO0001, Deaths From Political Violence, is coded (1) for high and (2) for low.
The other variable, VAROOOO2, is a Civil Rights ranking with (1) high civil rights
and (2) low civil rights. Deaths from Political Violence is the dependent variable.
Below the table are the measures of association generated in this run. What
measures presented in this chapter would be most appropriate for interpreting
this table? (Recall also that Yule’s Q is really a special case of gamma for a_
two-by-two table.)
Measuring Association in Contingency Tables » 387

VARO0001 * VARO0002 Cross-tabulation

VAROOO02

1.00 2.00 Total

VAROOO001 1.00 Count ~ 4 6 10


% within 36.4% 75.0% 52.6%
VAROO002

2.00 Count 7 2 9
% within 63.6% 25.0% 47.4%
VAROO002

Total Count 1 8 19
% within 100.0% 100.0% 100.0%
VAROOO002

Directional Measures
Value

Nominal by- Lambda Symmetric 294


Nominal VAROOO01 233
Dependent
VAROOO002 250
Dependent

Goodman and VAROQOOO1 146


Kruskal’s tau Dependent
VAROOO02 .146
Dependent

Uncertainty Symmetric 140


Coefficient VAROOO01 109
Dependent
VAROOO02 silt
Dependent

Ordinal by Somer’s d Symmetric ~.382


Ordinal VAROOO001 —.386
Dependent
VAROOO02 378
Dependent
388 @ STATISTICS FOR THE SOCIAL SCIENCES

Symmetric Measures

Value

Nominal by Nominal Phi


Cramer's V 362
Contingency 382
Coefficient 357
Ordinal by Ordinal Kendall’s tau-b —.382
Kendall’s tau-c 2/7
Gamma ~.680

-N of Valid Cases - 19

Exercise 11.3
Calculate Q and @ for the following table:

10 4

Exercise 11.4
From the same data set, Number of Protest Demonstrations (A) is dependent and
Civil Rights (B) is independent. Each variable is coded (1) for high and (2) for low,
__as before. The data are run using SAS. Confirm your calculations from Exercise 11.3
and select the most appropriate measures of association.

The FREQ Procedure


Table ofA by B

A B

Frequency
Expected :
Col Pet 1 2 Total

1 10 4 14
8.1053 5.8947
90.91 50.00

2 | 4 5
2.8947 2.1053
9.09 50.00
Total 11 8 19
Measuring Association in Contingency Tables » 389

Statistics for Table of A by B


Statistic Value

Contingency Coefficient 0.4169


Cramer's V : 0.4587
Gamma 0.8182
Kendall’s tau-b 0.4587
Stuart’s tau-c 0.3989
Somer’s dC |R 0.3750
Somer’s dR | C 0.4091
Pearson Correlation 0.4587
Spearman Correlation 0.4587
Lambda Asymmetric C | R 0.3750
Lambda Asymmetric R | C 0.0000
Lambda Symmetric 0.2308
Uncertainty Coefficient C | R 0.1588
Uncertainty Coefficient R | C 0.1876
Uncertainty Coefficient Symmetric 0.1720

Exercise 11.5
The first table presented in Exercise 6.2 is reproduced below as an SPSS table with
Political Rights (VAROO003) dependent and the Corruption Perceptions Index
(VARO0006) independent. Calculate gamma and both asymmetric lambdas for the
table. (Don’t forget to use frequencies, not percentages, to calculate these measures,
and remember that the totals are used to find lambda but not gamma.) Following the
table are edited portions of the output from the run. Use them to confirm that you
correctly calculated the measures of association.

VARO00003 * VAR00006 Cross-tabulation

VAROOOO6

1,00, 200 3.00 4.00 Total

VAROOO03 =1.00 Count 6 4 2 0 2


% within 100.0% 100.0% 66.7% 0% 0
VAROOO006

(Continued)
390 << STATISTICS FOR THE SOCIAL SCIENCES

(Continued)
VAROOO06

1,00 2.00 3.00 4.00 Total

2.00 Count 0 0 1 1 y,
% within 14.3% 3 0.0%

VAROO006

3.00. Gount 0 0 0 3 3
% within 0% 0% 0% 242.9%. 1520%
VAROOO006

4.00 Count 0 0 0 3 3
% within 0% 0% 0% 42.9% 15.0%
VAROOO06

Total Count 6 4 3 7 20
% within 100% 100% 100% 100% 100%
VAROOO006

Directional Measures

Value

Nominal by Lambda Symmetric 429


Nominal VARQO003 Dependent 2/5
VARO0006 Dependent 462
Goodman and VAROO003 Dependent 520
Kruskal’s tau VAROO006 Dependent 425
Ordinal by Somer’s d Symmetric .763
Ordinal VAROO0003 Dependent .690
VAROO006 Dependent 885

Symmetric Measures

Value
Nominal by Phi 1.020
Nominal Cramer's V 209
Contingency Coefficient 714
Ordinal by Ordinal Gamma 1.000
N of Valid Cases 20
Measuring Association in Contingency Tables » 391

Exercise 11.6
Below is the second table from Exercise 6.2, also reproduced as an SPSS table, with
Telephone Lines Per 100 People (VARO0005) dependent and Percentage of GDP
From Agriculture (VARQ0002) independent. Calculate gamma and both asymmet-
ric lambdas and compare your results to those in the output.

VAR00005 * VARO0002 Cross-tabulation

VAROOOO2

1.00 2.00 3.00 Total

VAROOO0S5 1.00 Count 0 1 6 Zz


% within 0% 143%. 75.0% 35.0%
VAROOO02
2.00 Count 0 4 2) 6
% within 0% Bye 250% 30.0%
VAROOO02
3.00 Count 5 2 0 7
% within 100.0% 28.6% 0% 35.0%
VAROOO02
Total Count 5 7 8 20
: % within 100.0% 100.0% 100.0% 100.0%
VAROOO02

Directional Measures

Value

Nominal by Lambda Symmetric .600


Nominal VARO0005 Dependent 615
VAROO002 Dependent 583
Goodman and VAROQO005 Dependent 474
Kruskal’s tau VAROO0002 Dependent 47
Ordinal by Somer’s d Symmetric —.780
Ordinal VAROO0005 Dependent —.786
VAROO0002 Dependent ~774

Symmetric Measures
Value

Nominal by Nominal Phi 961


Cramer's V .680
Contingency Coefficient 693
Ordinal by Ordinal _ Gamma —.963
N of Valid Cases 20
392 @ STANSTICSIFOR THE SOCIALSCIENCES

Exercise 11.7
Below is the table from Exercise 6.4, reproduced as an SAS table with Civil
Liberties (B) dependent and Political Rights (A) independent. Calculate gamma and
the most appropriate lambda for predicting the dependent variable. Compare your
results to those in the output below.

The FREQ Procedure


Table of B by A

B A

_ Frequency
Col Pet / y 3 4 Total

8 0 0 0 8
66.67 0.00 0.00 0.00
2 4 Z 6) 0 6
33.33 100.00 0.00 0.00
3 0 0 3 0 3
0.00 0.00 100.00 0.00
4 @) ie) 0 = 3
0.00 0.00 0.00 100.00
Total iZ Z 3 3 20

Statistics for Table of B by A


Statistic Value

Gamma 1.0000
Kendall’s tau-b 0.8486
Stuart’s tau-c 0.7267
Somer’s dC | R 0.7730
Somer’s dR | C 0.9316
Pearson Correlation 0.9378
Spearman Correlation 0.8774
Lambda Asymmetric C | R 0.7500
Lambda Asymmetric R |C 0.6667
Lambda Symmetric 0.7000
Uncertainty Coefficient C | R 0.8273
Uncertainty Coefficient R |C 0.7055
Uncertainty Coefficient Symmetric 0.7616
Sample Size = 20
Measuring Association in Contingency Tables j» 393

Exercise 11.8
Below is the table from Exercise 6.7, reproduced as an SAS table with the Corruption
Perception Index (D) dependent and GDP/Capita (A) independent. Calculate gamma
and the most appropriate lambda. The results are also in the output below.
~

The FREQ Procedure


Table of B by A
D A

Frequency
COMPCE 1 2 2 4 Total

1 4 2 0 O 6
57.14 66.67 0.00 0.00
2 5 / 0 0 4
42.86 33.33 0.00 0.00
3 0 0 3 0 3
0.00 0.00 50.00 0.00
4 0 0 3 4 7
0.00 0.00 50.00 100.00
foal 7 3 6 4 20

The FREQ Procedure


Statistics for Table of D by A
Statistic Value
Gamma 0.9016
- Kendall’s tau-b 0.7586
Stuart’s tau-c 0.7333
Somer’s d C| R ‘ 0.7586
Somer’s dR | C 0.7586
Pearson Correlation 0.8774
Spearman Correlation 0.8571
Lambda Asymmetric C | R 0.5385
Lambda Asymmetric R | C 0.4615
Lambda Symmetric 0.5000
Uncertainty Coefficient C | R 0.5937
Uncertainty Coefficient R |C 0.5937
Uncertainty Coefficient Symmetric 05937
Sample Size = 20
394 STATISTICS FOR THE SOCIAL SCIENCES

Exercise 11.9
An association matrix is presented below. The first three variables were introduced
in Chapter 6; the last two variables are Deaths From Political Violence per one
million population and Imposition of Political Sanctions by the Government per
one million population.’

Political
CDP Phone Agriculture Rights Deaths Sanctions

GDP 1.00
Phones 96 1.00
Agriculture —.92 —.91 1.00
Political Rights .64 .66 —.60 1.00
Deaths =79 ~/9 .66 —.68 1.00
Sanctions 22 —17 .08 12 76 1.00

1, Examine the gammas between GDP, Phones, and Agriculture. Then look at how
each of the three correlates with Political Rights, Deaths, and Sanctions. Is there
evidence that GDP, Phones, and Agriculture are all really measuring the same
underlying variable (known as a factor)? What would you name that factor?
2. Look at Deaths and Sanctions. Do they also appear to be measuring a common
factor?
Note: We will do more with this type of matrix in the exercises for Chapter 13.
LLL LEELE IOL LLL IION LLSDESEBEL LLL LED LL DOLE DESL EDDIE EEL LEER LES NOELLE RE LOL EON NE CTE RE NS OB ES tS

NOTES

1. It is also argued that for O to be valid, there must be at least 5 in each cell,
and the ratio of marginal totals should not exceed 30:70. If these criteria are not met,
find @ in lieu of Q.
2. L. A. Goodman and W. H. Kruskal, “Measures of Association for Cross-
Classifications,” Journal of the American Statistical Association, 49 (1954): 732-764.
3. A modified version of this table, used in earlier editions of this book,
was based on a recoding of data on a sample of countries based on information
contained in C. L. Taylor and D. Jodice, World Handbook of Political and Social
Indicators (New Haven, CT: Yale University Press, 1983).
—a ae

retracting amiss ru) ee


Seti, o—e=s =
ny
an ony
i
re & TL oe = oi ale

wien nal
seo
a,
utes — an «—@ou6
Ce, SE — in ee
nou

= -e INE
ih

a on edi 6° :
Field
_ are : >=
W KEY CONCEPTS ¥

chi-square test for observed frequencies/f, asymptotic standard


contingency/ expected frequencies/f, error (ASE)
contingency chi-square/ _Yates’s correction
chi-square test for for continuity
independence Fisher’s Exact Test
UBER SELLE EEE ODE EOS TESS OY JE:MCU NCU ES SUI EEE SESE EEE EINE YS NST USEE SEINE NE SOE IEEIR SNE ELSIE NEN ERIE SELENE LES
The Chi-Square Test

YW PROLOGUE ¥

The tests of ne er we eee eee up to eis point five


dealt mainly with comparing means and, less often, comparing variances
or variance estimates. Now we turn to the chi-square test for contingency.
This test is designed to test the statistical significance of a relationship in a
contingency table. Hence, the test’s name refers to contingency.
Note that the theme of contingency tables has been carried throughout
this book, in Chapters 1, 6, and 11, and now Chapter 12. We have moved
from describing data in tables, to calculating and interpreting percentages in
them, to measuring the association in them. Now, with a discussion of statis-
tical significance in tables, we complete the task first ae in Chapter 1.
WW \“—(6(:! WMP’. LSS_ EES RN % SWE DiD”™0UNDUIYJ =:XWEE.i WS. WW(QOSEEUL.-~XSS SLUM SEED EEE
ES YER SUP. CM SUM ll

pe 397
398 @ STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
In this chapter, we turn our attention back to tests of significance to examine
the chi-square test for contingency, a measure appropriate for cross-tabulations.
The variables in such tables may be any level of measurement—nominal,
ordinal, or interval—which is one reason why this test is so popular.
Although today we generally prefer to create indices of measurement that
approximate interval-level data and to which the previously covered tests of
significance would apply, there are still many instances where such index con-
struction is not feasible. Accordingly, we fall back on chi-square when nominal
or ordinal data must be analyzed. Thus, it is still useful for us today, but in the
early days of quantitative social research, its use was even more widespread.
There are many other tests of significance besides chi-square designed
for less than interval-level data. Some are even more powerful than
chi-square. However, surprisingly few of these other tests appear in the
research literature. This is partly because some of those tests are designed
for ungrouped rankings, individual ordinal data, which we rarely use. But
probably the major reasons for the popularity of the chi-square test are its
versatility and relative ease of calculation.

THE CONTEXT FOR THE CHI-SQUARE TEST

In Table 12.1, we reintroduce the death penalty/drug intervention prob-


lem presented in Chapter 11. The relationship from our random sample of
40 students at a specific college has already been confirmed (O = .80 and
@ = .50). How safe are we in generalizing that the relationship holds for all
students in the college from which our sample was drawn?
In effect, our null hypothesis postulates that within the entire
population of students in that college, there is no relationship between
support for intervention against foreign cocaine production and support
of the death penalty for domestic drug kingpins. That being the case, the
Q or @ found for the sample of 40 students must be the result of sampling
error, and ideally, if we had studied the entire population, O and @ should
both be zero, We thus could express H, in words as follows:

H,: In the population, the two sets of attitudes are unrelated.

We could also express H,, symbolically, selecting either of our two calculated
measures of association; for instance,

AL: ® Sopulation =0

Our H, could be likewise symbolically expressed.


The Chi-Square Test ® 399

A: Se oonlation #0

Note that our A, is nondirectional, in that it allows for either of two possi-
bilities: @,,, > 0 or @,., < 0. (Later on, we will work with directional alterna-
tive hypotheses for this test.)
~

Table 12.1

Attitude on Drug Intervention


Attitude on
Death Penalty For Against Total

Supports 5 5) 20
Opposes > 15 20

Total 20 20 40

If we can reject H,, then the relationship for our sample holds for the
entire population at the college, and the Q (or @) calculated from the
sample becomes a practical estimate of the relationship in the entire popu-
lation. We attempt to reject H, by means of a test of significance known as
the chi-square test for contingency (or the contingency chi-square,
the chi-square test for independence, or, most often simply, the chi-
square test). We often see this designated using the square of the Greek
letter chi as the symbol, y’, which, unfortunately, is difficult to distinguish
from a capital X*. Also note that chi is pronounced “kye.”

SLL SSOSSSAOLLLLULL LUTTE SSSSEER EES LESS MUTE


SEES TE BEBE SMES ELLE AULD ER AEE DME SSE RAE DME EL IES SULLA SES

Chi-square test for contingency A test of significance primarily used in a cross-


tabulation or contingency table.
SSRN AOA RSENS 8 LEDERER SEER ESSE SSI RELREE CME EDR
EAMES

Before we begin trying to reject our null hypothesis, we ask ourselves,


“What kind of a table would we have expected to get if H, were true?” The
cell entries in Table 12.1 are what we actually observed in our study. They
are called observed frequencies (or fs, where /stands for frequency,
and the subscript is the letter o, for observed). Now we ask, “Assuming
(a) that H,, is true and (b) that the marginal totals in the observed fre-
quency table actually reflect those marginals in the population (no. of
intervention supporters = no. of opposers, etc.), what are our expected
frequencies (fs)?” (You guessed it: f means frequency and e means
expected.)
400 << STATISTICS FOR THE SOCIAL SCIENCES

noe
ea oo oaeeeeeeammnneneeneenmmanmenenamneemmmaaaantll

Observed frequencies The actual frequencies observed in the contingency table.

Expected frequencies The frequencies we would expect to find if the null


hypothesis were correct.

Expected Frequencies

To find the expected frequency for any cell, multiply its row marginal
total by its column marginal total and divide by the grand total.

row marginal x column marginal


téc=
grand total

The expected frequencies are calculated below and presented in Table 12.2.

Table 12.2 Expected Frequencies

Attitude on Intervention
Attitude on the
Death Penalty For Against Total

Supports 10 10 20
Opposes 10 10 20
Total 20 20 40

Cell
20x20 400
Supports, For =——_ = =10
ne Jain ea 40
ZO S20. oHOU
Supports, Against les —— = — = 10
: 40 40
20x20 400
Opposes, For fe = ——_ = = 10
40 40
et SA 20. oe 400
Opposes, Against (= = = 10
40 40

Several things should be noted. First, the fs are all equal because the
marginal totals were all equal. This is rarely the case. Suppose we had the
marginals shown in Table 12.3.
The Chi-Square Test » 401

Table 12.3. Observed Frequencies

Attitude on Intervention
Attitude on the
Death Penalty For Against Total
Supports e milk 2 15
Opposes 5 20 25
Total 18 22 40

Then the fs would be

Cell
1 1 2
Supports, For Te 2 = i = — — oO
4
1 22 0
Supports, Against ie ~ = — = 6.25

2 18 450
Opposes, For Te “* = - a
+
25x 22 0
Opposes, Against ie ~~ = ~ = 13.75

Though logic would seem to suggest that we round these to whole numbers
(because we cannot have .75 or .25 of a person), from a mathematical
perspective, it is preferable to keep the decimal places. Thus, the expected
frequencies are as shown in Table 12.4.

Table 12.4 Expected Frequencies

Attitude on Intervention
Attitude on the
Death Penalty For Against Total

Supports 6.75 8.25 15.00


Opposes 25 ie 25.00
Total 18.00 22.00 40.00

Note that the marginal totals and grand total are the same in the f, table
as in the f, table. If they ever differ slightly, it will be due to rounding error,
as a consequence of having decimal places among the expected frequencies.
402 << STATISTICS FOR THE SOCIAL SCIENCES

Note also that while the


f{s are not all the same number, as they were in
Table 12.2, the rows and columns in an f, table are proportional to one
another. If we percentaged Table 12.4 by column, we would get the results
shown in Table 12.5.

Table 12.5. Expected Frequencies as Percentages of Column Totals

Attitude on Intervention
Attitude on the
Death Penalty For Against

Supports 37.5% 37.5%


Opposes 62.5 62:5

Total 100.0% 100.0%


=) (18) (22)

Table 12.6 Expected Frequencies as Percentages of Row Totals

Attitude on Intervention
Attitude on the
Death Penalty For Against Total i=)

Supports 45.0% 55.0 | 100.0% (15)


Opposes 45.0 Spe | 100.0% (25)

Note in Table 12.5 that the death penalty supporters are 37.5% of both
those for and those against intervention, and death penalty opponents are
62.5% of both groups. We would get similar results if we percentaged by row
totals as in Table 12.6.
If we interpret Tables 12.5 and 12.6 as we interpreted percentage tables
earlier in this text, we would conclude from the
fs that there is 70 relation-
ship between the variables. That is exactly what the null hypothesis states,
In fact, if we go back to the expected frequencies in Table 12.2 and Table 12.4
and calculate OQ or @ for each, they would be 0. Since earlier we stated our
null hypothesis in terms of @, let us calculate @ and see.
For Table 12.2,

7 ad —bc u (10)(10) — (10)(10)

eS Vatbetdyatob+d) J(20)(20)(20)(20)
= 100 — 100 « 0
= 0
~ ./160,000 400
The Chi-Square Test }» 403

For Table 12.4,

ed ad = 0c VOPR) = 6250025)
(at+bj(c+d)\(atob+d) J (15) (25) (18) (22)
022 280
= a)
/ 148,500 385.36

You can confirm that Q is also 0 for each of the tables.


A final point about the expected frequencies. In calculating eachf,, we
used the formula

row marginal x column marginal


te=
grand total

Applying this to the upper left (Supports, For) cell, we calculated

fay Ee el
==7.10)
rl) or eG

We did this in all four cells to get Table 12.2, but we really only needed to
generate one f,, We could have subtracted from row and column marginals
to get the other three expected frequencies.

10 20
20
201-20 |40

The upper row marginal of 20 minus the Supports, For f, of 10 gives us the
Supports, Against f,; 20 - 10 = 10. Now we have

Subtracting the Supports, For f, of 10 from its column marginal of 20 yields


the Opposes, Forf,,20 — 10 = 10. The same can be done in the Against cat-
egory, also getting 10. Thus,
404 << STATISTICS FOR THE SOCIAL SCIENCES

For the problem in Table 12.3, we also need only calculate one f,, Having
found out that the f, for Supports, For is

lo aals 270
Cd ae yao
40 40 ;

we could subtract from its row marginal of 15: 15 — 6.75 = 8.25.

15.00

18.00 22.00 | 40.00

Subtracting the 6.75 from its column marginal total, we get 18.00 — 6.75 = 11.25.

his|
15.00

25.00

18.00 22.00 | 40.00

We can get the last f{,(Opposes, Against) from either its row marginal
25.00 — 11.25 = 13.75 or its column marginal 22.00 — 8.25 = 13.75. Thus,

18.00 22.00 | 40.00

We have reconstructed Table 12.4.


Therefore, for a two-by-two table, we need only calculate onef,,and the
others can then be obtained via subtraction from the marginal totals. This
characteristic is used later in our actual significance test. The number off/s
needed before you can get the rest from the marginal totals is the test’s
degree(s) of freedom, symbolized df and must be calculated before we con-
clude our test of significance. Note the resemblance to the logic in the for-
mula used to explain degrees of freedom in Chapter 8, where the concept
was first discussed in depth. Obviously, our 2 xX 2 table examples have one
The Chi-Square Test » 405

degree of freedom, df= 1. For any size table, we may obtain the degrees of
freedom from the following formula:

df= (number of rows — 1) x (number of columns — 1)

fora2~x2 table, df= 2-D2-D=Ax)D=1

fora 2 x 3 table, df= (2-1) -1)=(1)(2) =2


for a 3 x 3 table, df= @ — 1) - 1) = (2)(2) = 4 and so on

Let us summarize our first problem.

@ = .50 o=0

df = (2 7 1@ =e 1) = (1d) z 1 Aly: @ opulation = 0

OBSERVED VERSUS EXPECTED FREQUENCIES

If the null hypothesis is true, we would expect that our observed frequen-
cies would be identical to the expected frequencies. However, as the result
of sampling error, we might often encounter some slight deviation.
Thus, if H, is true, we would expect this:

Js deviate + 1 unit from the fs.


406 << STATISTICS FOR THE SOCIAL SCIENCES

Sometimes, less often, we might even get this:

"eet f,s deviate +2 units from the fs.

The greater the deviation, the less it is likely to be the result of sampling
error. Though extremely rare, we could even get the following:

Je
f,s deviate +10 units from the fs.

In this case, even though H, is true, based on sampling error, we conclude


a perfect relationship between the variables.
What the chi-square test does is to look at the deviation between
each f, and its respective f,. These deviations are squared to get rid of
negative signs, just as we did with the standard deviation. Each squared
deviation is divided by its expected frequency, forming a kind of proportion
of the f, (although conceivably, each “proportion” might exceed one), and
the “proportions” are added together to yield a single chi-square value
for each table. The greater the deviation between fs and fs, the larger the
chi-square. Following are the chi-square values for these tables in order of
increasing f, versus f, deviation.

I, —f, Deviation Chi-Square


Table Per Cell Obtained

0) 0
The Chi-Square Test » 407

REY) 1.6

eS 10.0
(This is our
original problem.)

at U8) 40.0

We will turn shortly to the technique for calculating the chi-square


values. For now, note that when fs equal their respective fs (no deviation),
chi-square is 0. As f,s begin to deviate from their fs, chi-square begins to
increase. The larger the f, versus f, deviation, the larger chi-square gets.
However, as previously mentioned, the larger the f/s deviate from the fs,
and thus the larger the chi-square value, the smaller becomes the prob-
ability of drawing a random sample and getting that large a chi-square as
the result of sampling error. There is always some probability of drawing a
sample from a population where H, is true and getting a large chi-square.
At some point, however, we should feel safe enough to reject the null
hypothesis and conclude instead that our chi-square obtained was not the
result of sampling error but instead reflected the fact that the variables in
the population are indeed related.
As with our other tests of significance, we base our decision on
whether to reject the null hypothesis by comparing the obtained chi-
square to the critical value of chi-square at the .05 level. As we will see
shortly, based on a table of critical values of chi-square, for a nondirec-
tional H, at the .05 level with one degree of freedom (a 2 x 2 table), chi-
square critical is 3.84. If achi-square obtained from a 2 x 2 table equals or
exceeds 3.84, we can reject the null hypothesis. Returning to the tables
with increasingf, versus f, deviations presented earlier and comparing
the obtained chi-square to 3.84, we can reach a decision about statistical
significance.
408 4 STATISTICS FOR THE SOCIAL SCIENCES

f,-f, Deviation —Chi-Square


Table Per Cell Obtained Decision

0 0 Not Significant

zal 0.4 Not Significant

+2 iG Not Significant

ie
5a20 +5 10.0 Significant
15 |20 (This is our Gey
original problem.) exceeds 3.84.
BU ee ‘i : Relecusa..)

1) 40.0 Very Significant

USING THE TABLE OF CRITICAL VALUES OF CHI-SQUARE

In making our decision to reject the null hypothesis or not, we make


use of a table that provides the mathematically determined critical values
of chi-square, at selected probability levels, for a wide range of degrees
of freedom. See Table 12.7 and the appendices at the end of the book. In
the table, the degrees of freedom are listed down the column on the far
left side—not unlike the table for the ¢ test—and the tail areas, or significance
levels, are listed along the top row. For example, go to the line corresponding
The Chi-Square Test » 409

to the degrees of freedom for the table and look under the probability, the
significance level. Thus, at one degree of freedom, chi-square critical at
the .10 level is 2.71. Remember, we do not use that level to decide whether
or not to reject H, in most classroom situations (but we often use it in
nonacademic settings). For us, the crucial critical value of chi-square is the
one at the .05 level of significance found in Table 12.7.

Table 12.7 = Critical Values of Chi-Square

df 10 05 01 O01

1 2.71 3,84 6.64 10.83


2 4.60 5.99 9.21 13.82
3 6.25 761 11.34 16.27
4 7.72 9.49 13.28 18.47
5 9.24 107 15.09 20.52
6 10.64 12.59 16.81 22.46
7 12.02 14.07 18.48 24.32
8 13.36 1551 20.09 26.12
9 14.68 16.92 21.67 27.88
10 15.99 18.31 23.21 29.59
11 1728 19.68 24.72 51EZ0
ie 18.55 21.03 26.22 52.98
13 19.81 22.36 27.69 34.53
14 21.06 23.68 29.14 36.12
15 22.31 25.00 30.58 S70.
16 23.54 26.30 32.00 39.25
17 LATH 27.59 33.41 40.79
18 25.99 28.87 34.80 42.31
19 27-20 30.14 36.19 43.82
20: 28.41 31.41 37.57 45.32
21 29.62 32.67 38.93 46.80
22 30.81 33.92 40.29 48.27
23 32.01 2547 41.64 49.73
24 33.20 36.42 42.98 51.18
25 34.38 37.65 44 31 52.62
26 35.56 38.88 45.604 54.05
27 36.74 40.11 46.96 55.48
28 37.92 41.34 48.28 56.89
29 39.09 42.56 49.59 58.30
30 40.26 43,77 50.89 59.70
40 51.80 55.76 63.69 73.40
50 63.17 67.50 76.15 86.66
60 74.40 79.08 88.38 09:61
70 85.53 90.53 100.42 diz
SS TE res i Ss ee ie es pn ie
SOURCE: Abridged from R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricultural
and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint of Pearson
Education.
410 < STATISTICS FOR THE SOCIAL SCIENCES

If the obtained chi-square that we calculated is less than chi-square


critical at the .05 level, as in previous tests, we cannot reject the null
hypothesis. If the obtained chi-square equals or exceeds chi-square criti-
cal at the .05 level, we do reject the null hypothesis. The remainder ofthis
procedure also parallels those of other tests of significance. We compare
the chi-square obtained to the other critical values of chi-square appear-
ing in the table (.01 and .001 levels) and state our probability based on the
level of significance of the final critical value exceeded by the obtained
chi-square.
Earlier, we characterized a chi-square value of 40 obtained from a
2 x 2 (1 df) table as “very significant.” Now, note that at one degree of
freedom, the critical value of chi-square at the highest level of signifi-
cance in the table, the .001 level, is only 10.83. This is less than 40. Thus,
p < .001. In that same exercise, we are informed that for the original
problem (Table 12.1), we would obtain a chi-square of 10. We note the
critical values at one degree of freedom in Table 12.7. The obtained chi-
square of 10 exceeds 3.84, which is the critical value of chi-square at the
.05 level, so we reject the null hypothesis. At the .01 level, the critical
value is 6.64, which is also less than 10. However, 10 is less than 10.83,
the critical value at the .001 level. Accordingly, we report in this case only
thagp < OL
If we use a computer program to calculate chi-square, the exact
probability of alpha would be generated for us by the program. Below is
the information for another example as reported by SAS.

STATISTIC DF VALUE PROB

CHI-SQUARE 1 11.250 0.001

This tells us that from a one degree-of-freedom table, an obtained chi-square


value of 11.250 was generated. The probability of falsely rejecting a true null
hypothesis is exactly 0.001. (It is actually a bit lower, but it rounds up to .001
in this case.)
With this information, note that all we really have to do is examine the
probability generated. If that number is /ess than .05, then chi-square
obtained was greater than chi-square critical at the .05 level, and we may
reject the null hypothesis. If the probability on the printout was greater than
05, then chi-square obtained was less than chi-square critical at the .05 level,
and we cannot reject H,. Remember, the bigger the chi-square, the smaller
the probability, and vice versa.
The Chi-Square Test » 411

To summarize, for a specified df level:

If Decision Report
nel apes level X° not significant. No p statement
Do not reject H,
Xo ee lleva yh”. x is significant. p= 05
.05 level Reject H,
Xo DOL level > 7 ¥7_., 7 1s significant. p<.0l
.O1 level Reject. Hi,
ee Nee O01 level x is significant. p<.001
Reject H,

Before turning our attention to the actual calculation of a chi-square


value, let us briefly review the steps taken in running a test of significance in
general and the chi-square test in particular.

1. Before any data are examined, formulate the null hypothesis and the
alternative hypothesis.
2. Examine the data and calculate the appropriate test of significance.
For chi-square, this will require observed frequencies, expected fre-
quencies, the chi-square value itself, and the degrees of freedom.
3. Using a table of critical values, locate the appropriate critical values
for the degrees of freedom.
4. Compare the calculated test of significance—in this case, the
obtained chi-square value—to the critical value at the .05 level.
5. Ifthe obtained value is less than the critical value at the .05 level, we
cannot reject the null hypothesis.
6. If the obtained value exceeds the critical value at the .05 level, reject
the null hypothesis. If the obtained value is statistically significant,
examine the other critical values to determine what the probability
statement will be.
7. If the data are statistically significant, you may use these sample
statistics to estimate characteristics of the population (the popula-
tion parameters). Since for Table 12.1, we have been informed that
its chi-square value of 10.0 is statistically significant,p < .01, and we
may reject H,, we now conclude H,: In the population, attitudes
toward intervention against foreign drug producers and attitudes
concerning the death penalty for domestic drug kingpins are indeed
related. Thus, the @ = .50 value we determined from our sample is the
[Link] of what @ would be for our population of all students in
the college that we are studying.
412 @ STATISTICS FOR THE SOCIAL SCIENCES

CALCULATING THE CHI-SQUARE VALUE


Let us now see how to determine the obtained chi-square. We have already
formulated our hypotheses.

Ay: ® population —_

Hew population
#0

We have also been given the observed frequencies.

ip

Using the formula


row marginal x column marginal
grand total

and also using subtraction from marginal totals, we generated the fs.

Also, we have calculated the degrees of freedom.

df= (# rows — 1)(# columns — 1)


=(2-)2-)
= ()()
=]
Now it is time to see how the chi-square value of 10.0 that we were given
was actually calculated. The formula used was

2 (fo — fe)?

sage ee
We will use the following steps.
The Chi-Square Test ® 413

. Make a column labeled f, = and enter the observed frequencies


Starting with the upper left cell, moving across the top row until
completed. Then go down to the left side of the next row and con-
tinue across, and so on.
. Do the same thing with the fs, making sure eachf,is matched to its
respective f Thus,

i= f= |
15 10 |
5 10 |
5 10 |
15 10 |
The vertical line is optional; it is just a guide to where you begin
calculating the components of the chi-square value.
a Make a column labeledf,—f, = and subtract each /, from its respective
f, Ina 2 x 2 table (but only in a 2 x 2 table), these differences will be
either + or — the same value.
. Make a column labeled (f, — f,)? = and enter the square of the number
in the column generated in Step 3.
Gaon
nvake,a column labeled =, taking cach (ff) trom. the
column from Step 4 and setting up a division problem where the
divisor is that row’s /,.
GRSIDE mune
Do each &7 division.

Add up all the quotients from Step 6, thus producing ° oer which
is in fact chi-square.

ee 2 GSO:
eee eee a | ais
15 10. 4 5 25. 2510 = 25
5 TOs ot —5 2525 10625
5 10 | =5 25 25/10=— 25
15 10 | 5 se fis MOL ORS

y= Vo =fe)” = 100
fe

Recall that we then compare the obtained value to the critical values at
1 df and conclude that we can reject the null hypothesis, p < .O1.
414 € STATISTICS FOR THE SOCIAL SCIENCES

There is a special formula for calculating chi-square for a 2 Xx 2 table that


does not require the calculation of the fs. We do not present it here for two
reasons. First, it saves very little calculation time. Second, as we will see later,
we need to examine the f,s to make sure that chi-square is a valid test for the
data, and the special chi-square formula does not generate the needed
expected frequencies.
Note that the formula we are using here is based on actual frequencies
and not percentages.
Let us try calculating chi-square for the problem presented in Table 12.3.

P oi For Against Total


Supports 13 ys 15
Opposes 5 20 aS
18 22 40
Recall that we found for the upper left cell that

row marginal x column marginal


fe= grand total
_ 15x18
_270 _
ee (ee eae
Either by reapplying this formula to the other cells or by subtraction from
row or column marginals, we found the following expected frequencies:

ne
O75. d B25. 415,00

Las. 13.75) 25,00

18.00 22.00 | 40.00

Setting up and calculating chi-square:


. . fis a2 eae
Le ope Vitec] Ge nar ae
“=
113) Gos | 6.25 39.06 39.06/6.75= 5.78
2 825 eas 39.06 39.06/8.25= 4.73
Sy WS | —6.25 39.06 39.06/11.25= 3.47
ZOOM ow | 625 39.06 39.06/13.75 = 2.84

gap & aa = 2

te
df = (# rows — 1) x (# columns - 1) = (2-1)(2-1) = GD) Sal
The Chi-Square Test ® 415

Referring to the critical values of the chi-square table (Table 12.7) at


1 df
we find

? 2

x critical x obtained.

Xcriica OF = 3.84< 16.82 reject H,


eee Ol 6.64 < 16.82
Xe, O01 083i 16:82 sop 001

YATES’S CORRECTION
We noted earlier that chi-square increases from a table with no relationship
in it (chi-square = 0) to one with a perfect relationship (chi-square = 40). For
this table whose marginal totals were

20 20 | 40

we went from this table: ¥* = 0 to this table: y* = 40

fe

The f, —/, deviations went from 0 per cell to + 10 per cell. Given that set
of marginal totals, only 11 possible tables could be generated: f, —f, = 0, 1,
2,3,..., 10. Thus, only 11 chi-square values can be calculated for the given
marginal totals. Because the actual values of chi-square that can be gener-
ated in a low degrees-of-freedom table are so few, some statisticians have
argued that the chi-squares obtained only approximate the smooth curve
of the theoretical sampling distribution of chi-square. This could result in
generating an obtained chi-square large enough to reject the null hypothe-
sis, whereas, if the table had more cells, as shown below, chi-square
obtained would not have been large enough to reject the null hypothesis.
416 @ STATISTICS FOR THE SOCIAL SCIENCES

Attitude on Drug Intervention

Attitude on Strongly Moderately Moderately Strongly


Death Penalty Favorable Favorable Unsure Opposed Opposed

Supports
Opposes

For this reason, some—but by no means all—statisticians will adjust


the chi-square for a 1 df table by applying a factor that lowers the value
of chi-square obtained, making it harder to reject the null hypothesis. This
adjustment is known as Yates’s correction for continuity. (Some writers
call it Yates’s correction for discontinuity. Its the same thing, depending
on whether you see the glass as half full or half empty. We are correcting the
discontinuity to achieve the effects of continuity.) The correction takes the
absolute value of each f, — f, and subtracts half a point before going on to
square the number.

Yates’s correction for continuity —An adjustment of the chi-square for a 1 df table
by applying a factor that lowers the value of chi-square obtained, making it harder
to reject the null hypothesis.

Therefore, for our first problem where uncorrected chi-square was 10.0,
we use the following format to get the corrected chi-square.

ib i= 1 eee Ae eal | Sl
15° 1G) | 5 5 4.5 20.25
5 1004) —5 5 4.5 20.25
5 10 | =5 5 4.5 20.25
cia | 5 5 4.5 20.25

eee es ae
fe
20.25/10 = 2.025
20.25/10 = 2.025
20.25/10 = 2.025
20.25/10 = 2.025

2
=
‘s [lfo —fel — .5]?
] = 8.100
X corrected a f,
Je
The Chi-Square Test » 417

At one degree of freedom,

j
CO ees Sate l reject 7,
Va er aor Oat fi p<.0l
Veet elo On

In this case, even though when corrected for continuity, chi-square fell from
10.0 to 8.1, our original rejection of H, at the .01 level remains valid.
For the second problem, where uncorrected chi-square was 16.82,
corrected chi-square falls to 14.22 (do the calculations to confirm this). We
may still reject H, at the .001 level. The use of the corrected chi-square has
been the subject of much debate among statisticians, and in many instances,
it has been dropped entirely.’ Let your instructor provide further guidance
on its usage in your academic discipline. Because some do use it and com-
puter programs still generate it, for purposes of consistency, this text will
assume that the use of the chi-square with the Yates’s correction is appro-
priate for two-by-two tables. This will give you experience in determining
under what conditions some say the corrected chi-square should be used.

VALIDITY OF CHI-SQUARE

To be a valid test of significance, chi-square usually requires that most


expected frequencies be 5 or larger. This is always true for a two-by-two
table. If larger than a two-by-two table, a few exceptions are allowed as long
as (a) no f, is less than 1, and (b) no more than 20% of the f,s are less than
5. For most of the tables we encounter, this means the following:
Actual Whole Number
of Cells Where the fs
Can Be Greater Than
Table Size [Link] 20% offs | 1and Less Than 5
Dy Pepi 4 0.8 : 0
ASS 6 2 i
B63} 9 1.8 1
3x4 12 2.4 2,
4X4 16 ee 2)
4X5 20 4.0 4
iA} 25 5.0 5

Remember, these limitations apply to expected frequencies, not observed


frequencies.
418 @ STATISTICS FOR THE SOCIAL SCIENCES

When a computer program calculates a chi-square value, it will usually


issue a warning that chi-square may be invalid when too many /,s are below
the level of tolerance. (Unfortunately, these same programs often do not
show the fs, unless you request them in setting up your run.) If chi-square is
invalid for a two-by-two table, there is an alternative test known as Fisher’s
Exact Test, which may be used in place of chi-square. Fortunately, most
commercial chi-square computer programs also calculate Fisher’s Exact Test,
a real plus because Fisher’s Exact Test is not an easy one to calculate by hand.

Fisher’s Exact Test An alternative test for a two-by-two table when chi-square
is invalid.

For larger than two-by-two tables, we must locate the offending expected
frequencies and modify the table by collapsing or combining categories until
all fs satisfy the size criteria. Figure 12.1 provides a flowchart for this decision-
making process.
Suppose we are studying the relationship between religious affiliation
and socioeconomic status (SES) and our study yields the following observed
frequencies:

Pi Religion |
SES Protestant Catholic Jewish Other — | Total
High 25 10 5 0 | 40
Medium 20 10 5 5 | 40
Low 5 10 5 0 | 20
Total 50 30 15 5 | 100
For a 3 x 4 table, we can tolerate two fs that are less than 5 but greater
than 1. (Again note that we must examine expected, not observed, frequen-
cies. There are two f(s less than 5 above, but it does mot mean that thef‘s will
be less than 5.) We generate the expected frequencies to find the following:

if Religion |
SES Protestant Catholic Jewish Other | — Total
High 20 12 6 2 | 40
Medium 20 iz 6 2 | 40
Low 10 6 a 1 | 20
Total 50 30 is. 5: eS
The Chi-Square Test » 419

Figure 12.1 The Chi-Square Test for Contingency Flowchart

Is the table a 2x2 table?

Yes (1 df) No (larger than 1 df)

Are any f,’s less than 52 Are all f.’s grea’ than 12
7

Areno” = than 20% of


oe eess thanloie

yes No
od
yes No

|!
Chi-Square Do Chi-Square Do an
! !
Collapse or combine
is invalid with Yates Uncorrected rows and/or columns
correction for t until the above
continuity Chi-Square criteria are met,
! i.e.: All f.’s exceed 1,
;
Do Fisher’s dno more
Exact Test y y thana 20% 5 of the f,’s
,
are less than 5

ee
72 [Ifo aid f.| fa mile te) = ihe y
(f,

Refer to an Begin again at the


appropriate top of this flow
statistics text or chart
computer program

Remember that the above decisions involve examing the expected


frequencies, the f,’s, not the observed frequencies !!!

We see that there are four fs below 5: Low, Jewish and all three under
the other religions category. Not only that, but the Low, Other f, is 1. Thus,
chi-square is not valid as the table stands.
Note that most of the offending f/s are in the other religions category.
Because only 5 out of 100 people fell in that category, we could simply
delete the Other group from the study and collapse the table into a 3 x 3
table.
420 << STATISTICS FOR THE SOCIAL SCIENCES

e Religion
SES Protestant Catholic Jewish
High Z5 10 5
Medium 20 10 5
Low 5) 10 B,
Total 50 40 15
The expected frequencies are now the following:

is Religion
SES Protestant Catholic Jewish Total

High Zit 12.6 6.3 40.0


Medium 18.4 Pal ED 35.0
Low 10.5 6.3 3.2 20.0
Total 50.0 30.0 158) 5.0)

Now, all fs but one (Low, Jewish) are 5 or above, and the Low, Jewish f, is
greater than 1. Since in a 3 x 3 table, we can tolerate that one f,, chi-square is
a valid test for the modified table.
If we did not wish to delete the Other category, we could instead com-
bine Other with one of the remaining categories as follows:

(1) Other Catholic Jewish (Other being Protestant


plus the former Other)
or (2) Protestant Other Jewish (Other being Catholic
plus the former Other)
or (3) Protestant Catholic Other (Other being Jewish plus
the former Other)

Suppose we select the third alternative, perhaps because we want to com-


pare Protestants and Catholics to non-Christians. (This assumes that those ini-
tially indicating Other were non-Christians. One has to be a bit wary, however,
as both Mormons and Eastern Orthodox may have been included in the orig-
inal Other category. Let us assume that this is not the case here.) We then have

H Protestant Catholic Other | Total


High 25 5 | 40
Medium 20 10 | 40
Low 5 | 20
Total 50 20 |
The Chi-Square Test ® 421

fe Protestant Catholic Other | Total


High 20 12 8 | 40
Medium 20 12 8 | 40
Low 10 6 A | 20
Total 50 30 20 | 100

One f, is below 5 (Low, Other), but it is larger than 1, and a 3 x 3 table


can tolerate one exception. Thus, chi-square would be a valid test for this
table and is calculated below:

H,: In the population, there is no relationship between religion and SES.

H,: In the population, religion and SES are related.

(Note: We could also have worded H, and H, using a measure like lambda:
gs Lae = 0.)

oe (fo —fe)* _
ee aoe we nna) = a cae
(4

25 20 | 5 25) 25720 = 1,250


10 12 | —2 4 4/12 = 0.333
5 8 | —3 9 9/8 = 1.125
20 20 | 0 0 0/20 = 0.000
10 12 | =? 4 4/12 = 0.333
10 8 | 2 4 4/8 = 0.500
5 10 | -5 252 25/10=2.500
10 6 Cs 4 16 16/6 =2.666
5 4 | 1 1 1/4 = 0.250

C= ae = 8.957 = 8.96
(2

df = (# rows — 1) x (# columns — 1) = @-)DG-)N)=2@2)= 4

Examining Table 12.7, at df= 4,

Kova
critical?
05 = 9.49>8.96

We cannot reject Hy. Chi-square is not statistically significant. We cannot con-


clude a relationship in the population between religion and socioeconomic
Status.
422 @ STATISTICS FOR THE SOCIAL SCIENCES

DIRECTIONAL ALTERNATIVE HYPOTHESES

Suppose we are investigating the possibility of a relationship between one’s


level of social activism and one’s income. Although we could consider
activism to be a function of income, let’s reverse the order in this example
and explore the possibility that a higher income could be a reward for social
activism. We create a 3 x 3 table in which the independent variable, social
activism, is listed as a three-category ordinal variable (high, medium, low),
and the dependent variable, income, is also a three-category ordinal variable
(high, medium, low). Since we have an ordinal-by-ordinal table, we may use
gamma to form our hypotheses.

Aly: ospeition 7

yg aanpopulation #0

A random sample of 50 respondents yields the following observed


frequencies.

ie Social Activism

Income High Medium Low | Total

High 10 5 5 | 20
Medium 0 5 5 | 10
Low 5 10 5 | 20

Total is 20 its | 50

Calculating gamma from the observed frequencies, we obtain a value of


.18, suggesting a small relationship. If 7, is true, we must attribute the
gamma of .18 to sampling error. If H,, can be rejected, we may conclude that
in the population, there really is a relationship between one’s social activism
and one’s income, and we would estimate that population gamma with our
sample gamma of .18. We test significance using chi-square, The following
expected frequencies are generated.

pe High Medium Low | Total


High 6 8 6 | 20
Medium 3 4 3 | 10
Low 6 8 6 | 20
Total 15 20 15 | 50
The Chi-Square Test ® 423

(ignore the fact that in this example, too many Js are less than 5. This
example intentionally uses small frequencies.)

he f= fhe Uae Se e

10 6 4 16 16/6 = 2.666
5 8 3 9 9/8 = 1.125
5 6 =i 1 1/6 = 0.166
0 3 ~3 9 9/3 = 3.000
5 4 1 1 1/4 = 0.250
5 3 2 4 4/3, = 1,333
5 6 =i 1 1/6 = 0.166
10 8 2 4 4/8 = 0.500
5 6 el 1 1/6 = 0.166
x" = SS (fo Je) es 9.372

ie
af= (# rows — 1)(# columns — 1) = @ —- 1) - 1) = (2)(2) =4

Checking the table of critical values, we find

X’witicar -OD level, 4 df= 9.49 > 9.372


WesCa noe Teject A Xe eX acess
Under the rules we have set so far, we can go no further. We could,
however, redo the logic of the problem by making use of prior knowledge,
which might enable us to reject the null hypothesis. We go back to “square
one,” prior to collecting our data, when the null hypothesis was formed.

A): penueon =)

Now if we reject H, and conclude H,, we are saying Y,,putation * 9. Two possi-
bilities are subsumed under such an inequality:

either ¥,population >0 (gamma is positive)

Or

Notion <0 (gamma is negative)

Before looking at our data, we may have no reason to assume that one
or the other condition is more appropriate. Suppose, though, that based on
previous knowledge or experience, we have reason to believe that one of
those possibilities is logically impossible (or at least very improbable).
In this specific problem, suppose that based on previous research, there is
424 <q STATISTICS FOR THE SOCIAL SCIENCES

reason to believe that a positive gamma is likely and that a negative gamma
is highly unlikely. Thus, it is logical to assume that in the population, high
activism levels go with high incomes, and lower activism levels go with lower
incomes. In our actual sample, this will turn out to be the case with gamma
a positive .18, indicating clustering on the main diagonal.
Note, however, that if the reverse were true, we might have come up
with the table below.

ie Social Activism
Income High Medium Low | Total
High 5 5 10 | 20
Medium D) 5 0 | 10
Low 5 10 5 | 20
Total 15 20 15 | 50

Gamma for this table is —.18, but if we calculate chi-square, we will still get
a value of 9.372, the same as the one from the table where gamma was +. 18.
This means that for any given chi-square on the sampling distribution, half
of all the tables generating that chi-square value reflects tables where
gamma is positive, and the other half reflects tables where gamma is nega-
tive. That is, half of the area under the chi-square sampling distribution
reflects chi-squares from tables where gamma is positive, and the other half
reflects chi-squares from tables where gamma is negative.
Now if we may exclude in advance that either Y,.uiation >Oory population <0
is illogical, in effect we reduce by half all theoretically possible chi-square
values in our sampling distribution. Thus, the probability of obtaining any
specific value of chi-square doubles. This means that the probability levels
listed along the top of the table of critical values of chi-square may be cut
in half. At 4 degrees of freedom, what was

becomes

df 05" 1025" 005. «0005


4 7.78 9.49 13.28 18.47
Each of the first probabilities is divided by 2 to produce the second proba-
bility: .10/2 = .05, .05/2 = .025, and so on.
The Chi-Square Test 425

This being the case, chi-square critical at the .05 level drops from the
nondirectional 9.49 down to 7.78. Since the obtained chi-square value of
9.372 remains unchanged, our new conclusion is

X cue -0> level, Adf=7,/8 < 9.372 reject,’ p< :05


But A, is modified:

Ey population
> 0

By implication, though not stated, Y,,..ujaion < 9 is excluded in advance as


being illogical. In other situations, H, might have been

E,: ¥,population <0

implying that Y,,,,u1aion > 9 is, based on prior knowledge, illogical. Recall that
we followed the same logic with all our other tests of significance, with the
exception of f whose H, is always nondirectional.

TESTING SIGNIFICANCE OF ASSOCIATION MEASURES


In the previous examples, the chi-square test was used to determine statis-
tical significance. In several instances, it is possible to test for significance
by converting a measure of association into a z score and comparing the
obtained z to the critical values developed earlier in this text.
In the case of gamma, an approximation of z can be determined with
the following formula:

pe Ps +Pp

~ "Vand —y)

Recall that the religion versus SES problem yielded a chi-square of 8.96,
too low to reject H,. If, in the same table, religion had been replaced by an
ordinal variable such as the number of automobiles one owns—many, one
or two, none—then gamma would be an appropriate measure for the table.
If we calculated gamma for that table, it would be

—Pp es— eas


pias Ps+Pp
EP 1100 WFO
2450
Note:
n = 100
Rot P= 2450
y? = (.449)?= .202
1 —y?= 1 -.202 =.798
426 <4 STATISTICS FOR THE SOCIAL SCIENCES

Therefore,

Soe URINE ne9 ee en


n(l—y?) 100(.798)

Since 2 A9= 1.96, reject: p< .05.


A good rule of thumb, therefore, is to first do the chi-square test, and
if H,,is not rejected with chi-square, try the z conversion. In doing so, both
P, and P,, should exceed 100. An alternative for gamma and most other
measures of association is to calculate what is called an asymptotic
standard error or ASE. The conversion formula then becomes

the measure of association


= ASE

Most measures of association have accompanying formulas for finding


the ASE. Sometimes, these formulas are complex. Fortunately, a few statisti-
cal computer programs calculate the ASEs for us. SAS, for instance, lists ASEs
for gamma, lambda, and a number of other measures of association. They
appear to the right of the association measure in the printout. (These were
deleted in the exercises for Chapter 11 so as not to confuse you.) In SPSS,
the ASEs are similarly placed.

Asymptotic standard error (ASE) An alternative way of converting the


calculated measure of association to a z score.

Association Versus Significance

Association and significance are related but separate concepts.


Association measures show the strength of a relationship in a table, regard-
less of the size of the sample being studied. The significance level, by
contrast, increases with increasing sample size. If the sample size is large
enough, even tables whose association is small may be statistically signifi-
cant. Recall the discussion in Chapter 9 on the difference between statistical
significance and research significance. A conclusion about research signifi-
cance, relevance, must be based not on the size of the computed chi-square
value but on the size of phi or some alternative measure of association. To
illustrate this, let us calculate phi and chi-square for a series of tables all
having the same amount of association in them but having differing sample
The Chi-Square Test » 427

sizes. Since this is an illustration, we will use two-by-two tables and calculate
chi-squares without Yates’s correction for continuity. We will also ignore the
smallness of many of the expected frequencies.

Case j7 =

Sle 2a) eO)

_A2M-OMay_2-1_1_ 166
/B)Q2G)2) 736

a, (fo — ie
fe)”
Io= an i at ao i =
e

a 1.8 WA 04 47S = 0.022


J EY =e 04 .04/1.2 = 0.033
1 h.2 =A 04 04/0 2=0;055
1 0.8 VA .04 .04/0.8 = 0.050

y= 0186

df= (# rows — 1)(# columns — 1) = (2 — D2=1)S=QH@)=1

At 1 df 2critica = 3-84. Since 3.84 > 0.138, do not reject Hy. Chi-square is not
significant, and @ = .166 is quite low.

Case 2, 2 = 50

MOGmeOO
221 12 1
WAs6456
NOOO).
428 << STATISTICS FOR THE SOCIAL SCIENCES

fe f= fhe Gps 8 -
2

20 18 od 4 4/18 = 0.222
10 12 Jt 4 O12 = 555
10 12 =Z 4 4/12 = 0.333
10 8 Zz 4 4/8 = 0.500

ye 36s
Chi-square, though larger than in Case 1, is std// not significant; phi remains
unchanged from Case 1.

Gase 37 = 500

i 100 |300

300 200 | 500 300 200 | 500

oe_ (200)(100) — (100)(100) _ 20,000 — 10,000 _ 10,000


= = = .166
</(300) (200) (300) (200) +/3,600,000,000 60,000

ay
(eats it. ME) eeg ae 3

200 180 Z0 400 400/180= 2.222


100 120 =20) 400 400/120= 3.333
100 120 —20) 400 400/120= 3.333
100 80 20 400 400/80 = 5.000

¥2= 13.888
Not only can We reject FI; but since at 1 df, Peon nondirectional, .001 level a 10.83 a
13.888, p < .001. But while chi-square is significant and p < .001, the phi
remains unchanged at a low .166.
In contrast, in Case 4, the phi is larger.

Case 477 = 5
i | :

_ QQ=-MO _ 4-0 4
BAA) V3 6
The Chi-Square Test » 429

(iL ob= = ie ee
2

Se
2 2 8 64 .64/1.2 = 0.533
il 1.8 —8 64 .64/1.8 = 0.355
O 0.8 —8 64 .64/0.8 = 0.800
2, eZ ; 8 64 .64/1.2 = 0.533

(2= 2.221
Since chi-square critical is 3.84, the difference is not significant, even though
phi is relatively large.

CASE 572 = 50

_aS
(20)(20) — (10)(0) _
ce
400-0 400
e CEN = —— = 666
* ~ [G0 20)20)G0) 360,000 600
: (fo —fe)*
ee eee aes
e

20 12 8 64 64/12 = 0.533
10 18 —8 64 64/18= 3.555
O 8 —8 64 64/8 = 8.000
20 12 8 64 64/12 = 5.333

y2= 22.221
Not only is the chi-square statistically significant, butp < .001.
As discussed earlier, significance is a function of both the amount
of association and the sample size. So if you encounter a problem such
as that in Case 4 with high association but no statistical significance, an
increase in the magnitude of7 will yield the same association and statisti-
cal significance.

CHI-SQUARE AND PHI


In the previous chapter, we learned a formula for calculating phi for
a two-by-two table and learned that we use it only on a two-by-two table.
430 <4 STATISTICS FOR THE SOCIAL SCIENCES

In fact, however, phi may be calculated for a table of any size, by using the
following formula in which the letter 7 indicates the grand total for the
table.

tN

istX
n

While we may calculate phi and phi-square for any table, when the table
is larger than a two-by-two table, the maximum possible value of phi may
not be 1.0. It could be larger or smaller than 1.0. Thus, a given phi value is
difficult to interpret.
To correct for this, Pearson developed the contingency coefficient,
C, calculated as follows:

While the maximum possible C cannot exceed 1.0, it still can be less than 1.0
(in a 2 x 2 table, it cannot exceed .71). To compensate for that problem, two
other measures have been developed, Tschuprow’s T (rarely used today)
and Cramer’s V (also Cramer’s V or Craemer’s V).

Where m (meaning minimum)


Pat Ware is whichever is smallest: either
# rows — 1 or # columns — 1,

Contingency coefficient, Tschuprow’s T, and Cramer’s V_ Alternative measures to


compensate for possible misinterpretations of phi.

For a two-by-two table, @ = C = V For a larger than two-by-two table, the


three may not all be equal. For obvious reasons, ©, ¥ and C are known as
chi-square-based measures of association and are appropriate at any levels
of measurement. Though today, they are often supplanted by the measures
of association developed by Goodman and Kruskal (as well as by others),
such as lambda and gamma, these measures are still in use and are com-
monly found on printouts where tables and their accompanying chi-squares
are generated.
The Chi-Square Test » 431

COMPUTER APPLICATIONS
Before continuing here, go back and review the computer applications
section of Chapter 6 since this unit basically expands what was presented
earlier, adding the measures of association presented in Chapter 11 and the
tests of significance covered here. We discuss SPSS and SAS only.

SPSS

Recall that we began by typing in our data, and when this was done,
we clicked

Analyze
Descriptive Statistics
Crosstabs

This led us to a Dialog box, where we moved VARO0001 (GDP/Capita)


to the Rows box and VAROO002 (Percentage of GDP From Agriculture) to
the Columns box. At that time, we intentionally ignored the button labeled
statistics and went instead to the cel/s button. This time, we will visit the
former and revisit the latter. Press statistics.
In the new Dialog box, we first click in the box to the left of chi-square,
to indicate that we want that test run. Below that are the measures of asso-
ciation for nominal-by-nominal tables: C, phi and Y lambda, and the uncer-
tainty coefficient. To their right are the ordinal-by-ordinal measures: gamma,
Somer’s d, Kendall’s tau-b, and Kendall’s tau-c. You may click to run any or
all of these measures. We will run lambda and gamma since those are mea-
sures we learned to calculate earlier. (Lambda is for nominal-by-nominal
tables, but we can also use it for this ordinal-by-ordinal table.) Now click on
continue. We go back to the first Dialog box, where we then click on the
cells button.
Recall that before, in this Dialog box, we clicked on column percent-
ages, to include them in the table. We do so again. (Of course, you could
get the other percentages if you wanted.) This time, though, note the word
counts at the top of the box. Below it, already automatically called for is
observed, for observed frequencies. Below that is the word expected. Click
to the left of expected to show the expected frequencies in the table as well
as the observed frequencies and column percentages in each cell. Press con-
tinue, taking you back to the original Dialog box. Then press o& to get your
output, which is reproduced in Table 12.8.
432 << STATISTICS FOR THE SOCIAL SCIENCES

Table 12.8 Crosstabs SPSS

Cases

Valid Missing Total

N Percent N Percent IN Percent

VAROOOO1 * 20 100.0% 0 0% 20 100.0%


VAROOO002

VARO0001 * VAROOO002 Cross-tabulation

VAROO002

1.00 2.00 3.00 Total

VAROOOO1 1.00 Count 0 1 6 Fi


Expected Count 1.8 2.4 2.8 7.0
% within VAROO002 0% 14.3% 75.0% 35.0%

2.00 Count 0 2 1 3
Expected Count 8 I iPP 3.0
% within VAROO002 0% 28.6% 12.5% 15.0%

3.00 Count 1 4 1 6
Expected Count is ZA 2.4 6.0
% within VAROOO02 20.0% 57.1% 12.5% 30.0%

4.00 Count 4 0 0 t
Expected Count 1.0 1.4 1.6 4.0
% within VAROOOO2 80.0% .0% .0% 20.0%

Total Count 5) i 8 20
Expected Count Se, 7.0 8.0 20.0
% within VAROOOO2 100.0% 100.0% 100.0% 100.0%

Chi-Square Tests

Value df Asymp. Sig. (2-sided)

Pearson Chi-Square DEOu 6 001


Likelihood Ratio Pe 45 0) 6 001
Linear-by-Linear 12.916 ik .000
Assn.
N of Valid Cases 20

Case Processing Summary


a. Twelve cells (100.0%) have expected count less than 5. The minimum expected
count is .75.
The Chi-Square Test ® 433

Directional Measures

Asymp. Approx.
Value Std. Error’ Approx.T’ — Sig.

Nominal by Lambda Symmetric .600 “aley| 2.956 .003


Nominal VARO0001 Dependent —_.538 157 2.19) .006
VAROO002 Dependent —.667 .167 2,097 .007

Goodman & VAROO001 Dependent —.390 miley .OO1S


Kruskal’s tau VAROOQO0O2 Dependent —_.538 138 002°

a. Not assuming the null hypothesis.


b. Using the asymptotic standard error assuming the null hypothesis.
c. Based on chi-square approximation.

Symmetric Measures

Asymp. Approx.
Value Std. Error* Approx. T’ Sig.

Ordinal by Gamma —930 .067 —6.888 .000


Ordinal
N of Valid Cases 20

a. Not assuming the null hypothesis.

b. Using the asymptotic standard error assuming the null hypothesis.

SAS

As was done before, click

Solutions
Analysis
Analyst

Key in the data just as in Table 6.21. Then, as before, click

Statistics

Table Analysis

Highlight A and click on the Row box to move it. Similarly, move B into
the Column box, just as you did in Chapter 6.
Click the tables button, and in the new Dialog box, under frequency,
click to add the expected frequency. The observed frequency is already
434 STATISTICS FOR THE SOCIAL SCIENCES

indicated. Under percentages, the column percentage is already indicated.


(Add others if you want.) Click ok.
Back in the main Dialog box, click statistics. In that Dialog box, click to
add chi-square statistics and to add measures of association. (With SAS, you
get all of the measures.) Click ok, and once back in the main Dialog box,
click oR again. The output is reproduced in Table 12.9.

Table 12.9 Crosstabs SAS

The FREQ Procedure


Table of A by B

A B

Frequency
Expected
Col Pct 1 py G: Total

il 0 1 6 i
ieyS 2.45 20
0.00 14.29 75.00
2 0 2 1 2)
0.75 105 12
0.00 20.07 12.50
3 1 4 1 6
eS 2a 2.4
20.00 57.14 WEES,
4 4 0 0 4
1 1.4 1.6
80.00 0.00 0.00
Total 5 F 8 20

Statistics for Table of A by B

Statistic DF Value Prob

Chi-Square 6 22.6105 0.0009


Likelihood Ratio Chi-Square 6 23.2496 0.0007
Mantel-Haenszel Chi-Square LZ OL Sy 0.0003
Phi Coefficient 1.0633
Contingency Coefficient 0.7284
Cramer’s V 0.7518

WARNING: 100% of the cells have expected counts less than 5. Chi-Square may not
be a valid test.
The Chi-Square Test ®» 435

The FREQ Procedure


Statistics for Table of A by B

Statistic Value ASE

Gamma —0.9298 0.0668


Kendall’s tau-b . —0.7691 0.1016
Stuart’s tau-c —0.7950 0.1154
Somer’s d C|R =), 7310 0.1019
Somer’s d R|C —0.8092 0.1036
Pearson Correlation —0.8245 0.0869
Spearman Correlation —0.8207 0.0969
Lambda Asymmetric C|R 0.6667 0.1667
Lambda Asymmetric R|C 0.5385 0.1568
Lambda Symmetric 0.6000 0.1514
Uncertainty Coefficient C|R 05379 0.1301
Uncertainty Coefficient R|C 0.4354 0.1075
Uncertainty Coefficient Symmetric 0.4812 OTL

Sample Size = 20

CONCLUSION

The Limits of Statistical Significance


Although tests of significance are useful adjuncts to research, their
importance and results may be exaggerated. For one thing, whenever we
reject a null hypothesis, there is always the probability of a Type I or alpha
error: falsely rejecting a true null hypothesis. Also, these tests assume ran-
dom sampling, whereas the actual selection of the sample might have
resulted in some bias. Thus, no test of significance can be as useful as the
continued replication (repeating) of a study resulting in similar conclusions
each time the study is done.
For example, in this and earlier chapters, we have been looking at
the relationship between attitudes toward foreign drug intervention and
attitudes toward the death penalty. For a random sample of 40 students at
a selected college, we found a statistically significant chi-square and con-
cluded that. the relationship in the table (@ = .50) reflected a relationship
among all students at the college where the students were surveyed. If,
for some reason, that college is atypical of other such institutions, the
436 “ STATISTICS FOR THE SOCIAL SCIENCES

relationship will not hold elsewhere, and we would be wrong in concluding


that the variables are related. The only way to find out is to run the study at
other colleges or universities and see if the results are similar. At some point,
we make a leap of faith and assume that a finding is true everywhere in
the country, possibly everywhere on the continent, or even everywhere in
the world. If cultures were similar, we might be right; what’s true for Ohio
and New York could also be true for Ontario and Nova Scotia, or Scotland,
or New Zealand. Since cultures differ, however, differing results may easily
occur. For example, suppose we find in the United States or Canada that
among licensed physicians, the proportion of men predominates over that
of women. The same could be true in many other settings but not neces-
sarily everywhere: Historically, in the former Soviet Union, female physicians
have outnumbered male physicians.
In these kinds of problems, we are not dealing with random samples
of a larger population but rather with self-selecting groups. For the people
who attended the college we studied, the death penalty and intervention
attitudes are related. That’s all we can conclude. If, however, the 40 students
we had studied had been randomly selected from the population of all
the college students in the world, then our significant chi-square could allow
us to conclude that the variables were related among all college students
everywhere.

Chapter 12: Summary of Major Formulas

Finding an Expected Frequency

row marginal x column marginal


tls grand total

h Chi-Square (uncorrected)

Fae eee
fe)’ df = (# rows — 1)(# columns — 1)fe

Chi-Square (corrected)

wo seared
Pees = DS ee df = (# rows — 1)(# columns — 1)
The Chi-Square Test » 437

Chi-Square-Based Measures of Association

where m (meaning minimum)


is whichever is smallest: either
# rows — 1 or # columns — 1.

EXERCISES
‘Exercise 12.1
Social Status

Income High Medium Low |


“High uy 1 Oo | &
Medium — 1 2 0 | 3
Low 0 0 a
8 3) 9. | 20
Generate the expected Vecuenci Even though these are too small for a valid
chi-square, generate that measure anyway. If Hy is rejected, find @, C, and V.
(Can you see why @ is sometimes hard to interpret in tables greater than 2 x 2?)

: Exercise 12.2
_ Imagine that the same relationship in Exercise 12.1 had been determined for a
~ sample n= 200.
Social Status

Income High Medium low |


High 70 10- pb. | €
Medium 10 20 0 - 30
Low 0 0 0 =| 0
80 30 66 =| 200
438 << STATISTICS FOR THE SOCIAL SCIENCES

Find chi-square and, if significant, find @, C, and V. How do these four measures
differ from those in the previous exercise?

Exercise 12.3
Since the table in Exercise 12.1 has too many low expected frequencies for a valid
chi-square, reduce it to a 2 x 2 table by combining high and medium into a single
category for each variable. Generate the fs for this 2 x 2 table. Although techni-
cally, a chi-square is still invalid, generate chi-square uncorrected, and if H, is
rejected, find @, C, and V. Also calculate @ using the special formula for a 2 x 2
table presented in Chapter 11. Confirm that both formulas produce the same
results, except that the chi-square formula does not tell whether @ is positive or
inverse. Examine the table to get the appropriate sign for @.

Exercise 12.4
Now redo the chi-square for the 2 x 2 table in Exercise 12.3 using Yates’s
correction. How does chi-square change? Does it change enough to alter a
decision to reject H,? Note that even though H, may still be rejected, we use the
uncorrected chi-square to find measures of association. Thus, we do not calculate
those measures in this problem.

Exercise 12.5

GNP High Medium low Verylow |

High 0 0 + 8 iso
Medium 0 + 6 4 | 14
Low 6 6 2 0 badd
6 10 ‘le [ee | 40

Generate the expected frequencies and begin combining rows and/or columns
until chi-square is valid. (Hint: Look for a 3 x 2 table.) Then generate and interpret
chi-square, @, C, and V.

Exercise 12.6

High Medium — Low |

High 12 3 0 ras
Medium 6 6 3 | 15
Low 0 12 3 | 15
Very Low 0 0 15 | 15
18 21 21 Fo 6s
The Chi-Square Test » 439

Generate expected frequencies. If necessary, combine rows until chi-square is


valid. Calculate chi-square and its measures of association. Interpret the results.

Exercise 12.7
S

High Medium Low Very Low |


High 2 3 0 0 eee
Medium 1 3 2 2 | 8
Low 0 1 2 3 oieG
3 7 4 5 eo
Generate expected frequencies. If necessary, combine or delete rows or columns.
Determine the appropriate test of significance.

Exercise 12.8
Calculate chi-square for the following table.

Exercise 12.9
Both variables have both been dichotomized in the table below.

Protest Civil Rights


Demonstrations High Low
High 10 4 14
Low 1 Ge co
1] 8 19

What is the most appropriate test of significance for these data? Calculate
both the uncorrected and corrected chi-squares. What are your conclusions?
Calculate @.
440 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 12.10
Although an uncorrected chi-square would not be valid on this table (and even if
valid would not be significant), calculate it anyway. Calculate . Compare this @ to
the one in Exercise 12.9. Remember to inspect each table to see if @ is positive or
negative. What are your conclusions? Explain them.

High Low |
High 0 2 | 2
Low 1 3 | 4
1 5 | 6

NOTE

1. J. T. E. Richardson, “The Analysis of 2 x 1 and 2 x 2 Contingency Tables: An


Historical Review,” Statistical Methods in Medical Research 3, no. 2 (1994): 107-33.
My thanks to Professor Richardson for sending me the article.
»
-_

IOC! Ari
nm Analysi

\aSH) 6 .« Ones
boer ® MOH | (bse . eG
Bass USAT) VF tae
; hula A 4 uf Vilasvy
Von :
ted
Gl es
) U oF

* Ai
Pa a ns Pity @ & y 6 :

is i
Sa a UA ® AohGe omer SDs oapee }
Rt hee 8S 64pede unre dy dag TS »
Ht fa), eyeliltnn ae ai? Inet ae

—, Var 1 eiyaeiiit e? : Aue Co


Se eye “ye rye eee 7 (
VW KEY CONCEPTS ¥
a NAN NI Mi

Pearson's r/r/the x-coordinate scatter diagram/


Pearsonian product- y-coordinate scattergram/scatter
moment correlation function plot
coefficient linearity least squares method
coefficient of linear equation regression ofy on
determination/r* constant x/regression of x on y
regression equation variable coefficient of alienation/
correlation-regression y-intercept t=
analysis slope y—the predicted
Cartesian coordinates positive versus inverse value of
origin of a graph (negative) relationship intraclass correlation
X-axis curvilinear versus linear coefficient (7, or 7,)
Y-axis relationship correlation ratio [EF or
ordered pair linear regression nN (eta)|
ROBE SSE LO
RN 22
CHAPTER

Correlation and
Regression Analysis

VY PROLOGUE ¥

Correlation and regression analysis, presented in this chapter and the


next, bring us back to the consideration of the strength of a relationship
between variables. This was covered for cross-tabs by our study of measures
of association presented in Chapter 11. Now, instead of data in tables, we
have individual interval- or ratio-level data for two variables. What is the
strength of the relationship between them?
Imagine that for each person in a sample, we know that individual’s
income and also have a score on a social status scale ranging from a low of 0
to a high of 10. How and how strongly are income and social status related?
Suppose we find that there is a relationship. Then we also might want
to develop a way to make predictions about others. Suppose I know your
income. What do I predict your social status score would be on the same
scale?
LH WC YJJH$ WS DMD. WS 2 SI/IBIENDNN. EEK Ns QIGEERBIYIEN

» 443
444 << STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
Correlation-regression analysis is a set of techniques that have gained
widespread popularity throughout the social sciences and in business and
economics as well. It is a set of interrelated techniques designed for indi-
vidual raw score data at the interval or ratio level of measurement. These
techniques are quite versatile and relatively sophisticated.
We first generate a correlation coefficient, a measure designed to
ascertain the strength of a relationship between two variables. (The mea-
sures of association presented in previous chapters were, in fact, designed
to do for data in tables what their forerunner, the correlation coefficient,
does for interval- or higher-level data.)
If the correlation coefficient thus generated is large enough to demon-
strate a meaningful relationship between the variables, we may then gener-
ate a linear regression equation. This is a mathematical equation designed
to predict what a subject’s score would be along the dependent variable if
we had knowledge of the subject’s score on the independent variable. Thus,
we go a step beyond measuring the strength of a relationship into the realm
of making predictions. The larger the correlation coefficient, the more accu-
rate our predictions will be.

THE SETTING

In Chapters 6 and 11, we examined relationships between grouped nominal


or ordinal variables presented in cross-tabular form. We also measured asso-
ciation by means of several techniques. Most tables in this text present data
at the nominal or ordinal level of measurement. From time to time, grouped
interval-level variables may appear in tabular format, but in general, the
grouping is done by the researcher for purposes of presenting the findings
to the reader. Originally, the data were individual, interval-level raw scores.
In such a situation, we have at our disposal techniques of analysis that are
more sophisticated and more useful than many of the techniques used to
interpret cross-tabulations. Known collectively as correlation-regression
analysis, these new techniques enable us not only to measure association
but also to make predictions of scores on the dependent variable. As
researchers, we ideally want to collect as much data as possible at the inter-
val level because we can make use of these more sophisticated techniques
of analysis to learn as much as possible about the variables under study. In
this chapter, therefore, we will study relationships between interval-level
variables that are not in tables but in ungrouped, individual raw score for-
mat. See Box 13.1 for a summary.
Correlation and Regression Analysis ® 445

BOX 13.1

When Is Correlation-Regression Analysis Used?

Two individual (raw score) variables measured by interval


scales.
~

Measure the strength of the relationship between these two


variables.
If that relationship is sufficiently strong, describe the nature of
the relationship between the two variables in such a way that
it will be possible to predict a respondent’s score on one vari-
able if we know that person’s score on the other variable.

Suppose we are doing a study of alleged political corruption. We have


reason to believe that a locally elected district attorney is demanding kick-
backs to his reelection campaign fund from those appointed to patronage
positions in his department. He, of course, denies this and claims that
any contributions are strictly voluntary, reflect no coercion or pattern, and
certainly do not reflect a specific percentage of the employees’ salaries.
From public sources, we acquire the names, salaries, and mean monthly
contributions made by each employee to his or her boss’s war chest.
Following are the findings, with the employees’ names discreetly replaced
with numbers.

x= Ve
Annual Income Mean Monthly
Employee (in thousands of dollars) Contribution
1 $80 $160
2 70 yD)
3 a2 97
4 45 85

Each of the four people is measured along two variables by individual, interval-
level raw scores.
Our questions are as follows: (a) Is there a relationship between one’s
income and one’s monthly campaign contribution? (b) If there is, how
strong is that relationship? (c) Is it possible to estimate (predict) someone’s
contribution if we know that person’s income and base that prediction on
the data for the original four individuals? (d) What is that prediction for the
person in question?
446 < STATISTICS FOR THE SOCIAL SCIENCES

Our first task is to determine if a relationship exists and, if so, how large
it is. To do this, we calculate a measure known as a correlation coefficient.
This measure resembles the measures of association for cross-tabs such as
Q, gamma, and lambda discussed in previous chapters. When data are not
grouped in tables but are individual raw scores, we refer to the measure of
the strength of that relationship as a correlation coefficient rather than a
measure of association. When both variables are an interval or a ratio level
of measurement, as is the case here, we use a coefficient formally known as
the Pearsonian product-moment correlation coefficient. Since that is
quite a mouthful, we usually refer to it as Pearson’s r or just r.

Correlation coefficient Measure of strength of a relationship in which data are not


grouped in tables but are individual raw scores.

Pearsonian product-moment correlation coefficient or Pearson’s r Coefficient that


is used when both variables are an interval or a ratio level of measurement.

Like gamma, Pearson’s 7 is actually measuring how linear the relation-


ship is (more on that later), and it ranges from 0 to 1 in value. It also carries
a sign—positive (for a positive relationship) or negative (for an inverse rela-
tionship). So r reads like a gamma. Once we know 7; we square that value to
obtain what we call the coefficient of determination, another indicator
of the strength of the relationship being studied. The coefficient of deter-
mination, 7*, will tell us the proportion of variation in the dependent vari-
able (j’) that can be explained by variation in the independent variable (x).
pissin

Coefficient of determination Indicates the proportion of variation in the dependent


variable (y) that can be explained by variation in the independent variable (x).
eames ae S808

If y and 7 are low, we may conclude that going any further is useless.
Even though we could develop a predictive model, it would not be precise
enough a predictor to be of value. If7 and 7 are large enough to interest us
(and this is for now strictly a personal judgment call), we may conclude that
it would be useful for us to be able to predicty from x—in this case, to pre-
dict someone’s campaign contribution from that individual’s income. That
predictive model, the mechanism for estimating ay score from the respective
x score, is known as a regression equation. Since these two steps are so
closely interrelated, we could call the procedure in its entirety correlation-
regression analysis. In fact, we do so in this chapter. However, it should
be pointed out that in actual research, the two procedures are often sepa-
rated, so sometimes you will see correlations but no regressions, sometimes
regressions but no correlations, and sometimes both correlations and
Correlation and Regression Analysis j 447

regressions together. Because the concepts of correlation and regression


are so Closely related, they are presented here as parts of a single technique,
even though in practice they may appear separately.
SLATE SN
eeRARReRANNRER SLES OSES SSSA LIS
ALLEN

Regression equation The mechanism for estimating a y score from the


respective x score.

Correlation-regression analysis The presentation of correlation and regression


techniques together.
cesonenssce essen atin cece TY

Before reaching the point of actually calculating and applying correla-


tion coefficients and regression formulas, we need to discuss a number of
preliminary topics. These topics—Cartesian coordinates, the concept of lin-
earity, and linear equations—may be old material to many readers who have
appropriate mathematical backgrounds. If so, treat these sections as neces-
sary review material. However, do not worry if these topics are new to you;
they are presented in enough detail for beginners. All readers should keep in
mind the three topics that follow are preparations for correlation-regression
analysis—not the analysis itself. The latter will follow in due course.

CARTESIAN COORDINATES

In analyzing data, a useful place to begin is with a pictorial representation of


the relationship between the two variables under study. We can do this by
means of Cartesian coordinates, a method of graphing relationships devised
by the French mathematician René Descartes. Not only is this method of
pictorial representation relatively simple, but it is conceptually elegant and
mathematically valuable. From it stems the field of mathematics known as
analytic geometry, which links curves on a graph to algebraic equations.
Many readers, in fact, will already be familiar with Cartesian coordinates.
We begin by constructing a graph composed of two perpendicular lines
called axes (singular: axis; plural: axes) (see Figure 13.1). Each axis is mea-
sured in scales resembling those on a ruler, with the point where the two
axes intercept given a value of zero. That point is the origin of the graph.
The axis that extends horizontally is the x-axis, and the one that extends
vertically is the y-axis. On both axes, the measurement numbers increase in
magnitude as one moves out from the origin. However, depending on the
direction one is moving, we give to each measurement number a + or — sign.
On the x-axis, the numbers to the right of the origin are positive (+) (we
usually do not put + in front of each number), and those to the left are neg-
ative (—). On the y-axis, the numbers above the origin are positive, and those
448 STATISTICS FOR THE SOCIAL SCIENCES

below it are negative. On each axis, the distance between each unit is the
same. For instance, if we lay a ruler on the x-axis, we see that the distance
from +1 to +2 is the same as the distance from +2 to +3 or from —4 to —5,
and so on. Note, however, that it is not necessary that both axes be mea-
sured in units of the same size. That is, the distance between +1 and +2 on
the x-axis need not be the same as the distance from + 1 to +2 on the y-axis,
as long as the distances between all unit intervals are the same on edch axis.

Origin The point of a graph where the two axes intersect indicating a value
of zero on each axis.

x-axis The axis that extends horizontally.

y-axis The axis that extends vertically.

In Figure 13.2, for instance, units of 1 on the x-axis are as far apart from
each other as are units of 10 on the y-axis (measure it yourself). The arrow-
heads at the ends of the axes are essentially reminders that the axes them-
selves are limited in length by the size of the pages on which they are
printed. Mathematically, the page on which the graph is printed delineates

Figure 13.1

y-axis

l
or
Correlation and Regression Analysis j» 449

Figure 13.2

y-axis

50

40
30

20

10
O ;
SS ln lO a OO Oe a), Sr
5 -4 -3 -2 4 eee 1 2 She Ave a5
=O

—20

=A) +

a plane, a surface in space that can be extended infinitely beyond the edges
of the page itself. Consequently, while in fact we run out of space in any real
page, conceptually our axes continue out indefinitely. (Want to impress your
friends? Say each axis approaches positive and negative infinity as its limit.)
Given the above format, we note that it is possible to numerically
designate any point on the graph’s page (or its logical extension beyond the
page) by relating that point to the two axes. Note the point we have desig-
nated P, in Figure 13.3. We begin by drawing two dotted lines, which go
from P, to the x-axis and y-axis in such a way that each dotted line is per-
pendicular to one axis and parallel to the other. The line /,, which drops
down from P, to the x-axis, is parallel to the y-axis and perpendicular to the
x-axis (the angle between /, and the x-axis is 90 degrees). The line /, is par-
allel to the x-axis and perpendicular to the y-axis. In the case of P,, we see
that /, intersects the x-axis at + 4 and /, intersects the y-axis at +3. We use
these two values to locate P,. By convention, we designate the point by what
is called an ordered pair, a set of two numbers in parentheses separated
by acomma. The first number is the x-coordinate, and the second number
is the y-coordinate. The x-coordinate is the number on the x-axis where /,
crosses it. The y-coordinate is the number on the y-axis where /, crosses it.
P, is thus designated by the ordered pair (+4, +3) or, more simply, (4, 3)
450 < STATISTICS FOR THE SOCIAL SCIENCES

Figure 13.3

X-axis

since we may assume each number to be positive unless specifically desig-


nated as negative.
REAR

Ordered pair A set of two numbers in parentheses separated by a comma,


indicating a point on a graph.

x-coordinate The first number in an ordered pair.

y-coordinate The second number in an ordered pair.

We may now find any point on the graph if we know its coordinates.
Suppose we are asked to locate the point P,, whose coordinates (x,, y,) are
given as (—2, 3), respectively (see Figure 13.4). Since the first number in our
ordered pair is always the x-coordinate, we locate —2 on the x-axis. Then we
locate the y-coordinate (+3) on the y-axis. From each of these points on the
axes, we construct our dotted lines (/, and /,) perpendicular to each of the
axes, and where the dotted lines intersect, we have our point P,.
The origin of the graph will always have the coordinates (0, 0). Note that
any point exactly on the x-axis will have a y-coordinate of 0, and any point
exactly on the y-axis will have an x-coordinate of 0.
Correlation and Regression Analysis >» 451

Figure 13.4

0
<o e > X-axis
5 -4 -3 -2 VA be 2 3 A 5
—1

Sl

if
3
54

THE CONCEPT OF LINEARITY

We use the Cartesian coordinate system to plot the relationships between


variables we would like to explore. By convention, we label the variable we
presume to be the independent variable as x and the dependent variable,
the one whose variation we are trying to explain, as y. Remember that in
research, we begin with a number of cases or respondents and assign to
each one a score or number related to each of the variables we are studying.
Thus, each respondent has both an x score and a y score. To be more spe-
cific, suppose that for a sample of six people, we wish to explore the rela-
tionship between two characteristics that we know each of the respondents
possesses. For each respondent, we can measure both characteristics with
at least an interval scale. We assume that one of the characteristics (the
dependent variable) is related to or can be explained by the extent to which
the respondent possesses the other characteristic (the independent vari-
able). Mathematically, we say that y is a function of x. In this case, let us
suppose that y, the dependent variable, is total savings and that x, the inde-
pendent variable, is the respondent’s level of education as measured by
number of years of school he or she has completed. The hypothesis we wish
to examine is that the more schooling one has, the more money one will
452 << STATISTICS FOR THE SOCIAL SCIENCES

save; savings is a function of education. If it is indeed true for everyone, we


assume that it will be true for our group of six respondents. Consequently,
we find out from each respondent his or her years of education (x) and
amount of savings (v) and present the data in tabular form (see Table 13.1).

Function The case where a score on the dependent variable (y) may be predicted
from a score on the independent variable (x). The value of y is obtained either
graphically or by an equation.

Table 13.1

x = Education y = Total Savings


Respondents (years of schooling) (in thousands of dollars)

Tom 0
Carol 4 2
Jim 8 4
Lucy 12 6
Jack 16 8
Jill 20 10

Notice that we can make use of Cartesian coordinates to graph each


respondent in terms of scores. If we construct an x-axis to measure educa-
tion and a y-axis to measure savings, and for each respondent we have both
an x and ay score, then each score may be treated as a coordinate in an
ordered pair. Thus, each person may be indicated by a point on the graph
as determined by that person’s education and savings (see Figure 13.5).
From the table and the graph, we immediately notice two things. First,
in every one of the six cases, the greater the person’s education, the greater
the savings. On the basis of our hypothetical data at least, our hypothesis is
confirmed. Second, we notice that the pattern of the points in Figure 13.5 is
an exact straight line. We confirm the fact by laying a ruler over the points
and drawing a straight line through them all, as we have done in Figure 13.6.
Since the relationship is an exact straight line, we say that the two variables,
education and savings, are linearly related. The straight line in Figure 13.6
is an exact visual representation of that relationship.

—ennatnttntetttmtttettttnetitnttnten
Linearly related Relationship that is shown as an exact straight line.
ene
Correlation and Regression Analysis » 453

Figure 13.5

12

>
wn
0
1 in
@ Jill
'
(20,10)
=
by a fal @ Jack (16,8)
ces
~ 64 @ Lucy (12,6)

4 | @ Jim (8,4)

2 ®@ Carol (4,2)
Tom (0,0)
= a Ss =| (tera Sia =a) omit a a i aaa ar | T >

0 2 4 6 S O Ww WW 16 ie we
x = Education
(Years of Schooling)

The straight line in Figure 13.6 enables us to predict a person’s score on


one variable if we know that person’s score on the other variable. If we
assume that the straight line holds for everyone in the population and
another person comes along with a specific number of years of schooling,
we can use the line to predict that person’s savings.
For example, suppose Ann informs us that she has had a total of 24 years
of schooling. To predict her savings, we make use of the fact that if the
straight line in Figure 13.6 reflects the relationship between education and
income for everyone, then the ordered pair representing Ann’s education
and income must also fall on that line. We know that Ann’s x-coordinate is
24, so if we designate her income as y, (the subscript a stands for Ann), then
the point (24, y,) will fall on that straight line. We locate Ann’s education
(24) on the x-axis and draw a perpendicular dotted line (/,) up from that
point until it intersects the line (see Figure 13.7). This will be the point (24,
y,). To find y, by visual inspection, we now construct a dotted line /, that
goes through (24, y,) and is parallel to the x-axis (perpendicular to the
y-axis). The line /, crosses the y-axis at exactly y,, at the spot where y = 12.
Thus,y, = 12. We predict, therefore, that Ann’s savings will be $12,000.
Now, suppose yet another person, Frank, tells us he has completed
10 years of school. To predict Frank’s savings, we find 10 on the x-axis and
erect a perpendicular /,, which crosses the line at the point (10, y,). Where
454 STATISTICS FOR THE SOCIAL SCIENCES

Figure 13.6

Y=

Figure 13.7
Correlation and Regression Analysis » 455

/, crosses the line, we draw a new line J, parallel to the x-axis and see where
/, crosses the y-axis. Since /, crosses the y-axis halfway between the 4 and
the 6, we conclude that y, is 5. We predict that Frank has saved $5,000.

LINEAR EQUATIONS
The predictive process we have just been using may also be accomplished
by means of simple algebra. If it is indeed true that income is linearly related
to education as the data in Table 13.1 would suggest, then it is possible to
represent the straight line used in Figures 13.5 to 13.7 by an algebraic equa-
tion. Specifically, that equation will be of a form known as a linear equa-
tion, affirming that the points generated by the equation will graph as a
straight line rather than any other kind of graphic figure.

Linear equation Equation in which the points generated will graph as a straight
line rather than any other kind of graphic figure.

By way of review, we should recall what an equation really does.


Essentially, an equation is a set of mathematical instructions—a program—
that matches specified values of some variable x with specified values of
some other variable y (it “maps” values of x into values of y). The equation
tells us how to find the particulary value that matches it with the value of x
that we have designated. We supply the value for x; the equation tells us the
corresponding value of y. The line in Figures 13.5, 13.6, and 13.7 may be
expressed algebraically by a linear equation so that instead of finding y
graphically, we can calculate it by solving the equation fory. All linear equa-
tions can be written in the following form:

y=a tox

In this equation, a and b are known as constants.' This is because each


straight line that we can draw on the graph will have a particular a-value and
a particular b-value associated with it, and for that specific line, a and b
remain the same—constant—no matter what value of x and y may be
plugged into that equation. By contrast, the x-value and the y-value are
called variables because any given straight line will have many different
values of x andy associated with it; these values vary from case to case. The
constants, a and b, determine where on the graph the line is located and
differentiate that line from all the others. The variables, x and y, relate to
specific points on that line.
456 << STATISTICS FOR THE SOCIAL SCIENCES

Constants Particular values of a specific linear equation that remain constant.

Variables Particular values of a specific linear equation that vary from person to
person (or unit of analysis to unit of analysis).

We are already familiar with the concept of a variable. Constants remain


constant for a particular line (a particular relationship), although @ and b do
indeed vary from one line to the next. Suppose for a given relationship we
know the constants a4 and b. We are left with x and y, the variables. But if x
is specified, we then know a, 6, and x and can easily solve fory, the remain-
ing unknown. Solving for y—predictingy for a given x—is the major goal of
this process.
The constants are not merely numbers drawn like rabbits from a magi-
cian’s hat; both @ and b can be defined in terms of the characteristics
of the particular line under study. The first of these, a, is known as the
y-intercept of the line. It is the value of y at that point where the line
crosses the y-axis. Note that because we have indicated that the axes
extend out indefinitely, we can specify that in every case other than where
a line is parallel to or on the y-axis, a line will eventually cross the y-axis
and thus will have a y-intercept. This can be seen by placing a ruler’s edge
on a graph, moving it and rotating it around. The line it delineates will
either cross the y-axis on the graph itself or on some extension of the line
and the y-axis (see Figure 13.8). The y-intercept is the value of y where
x = 0. Note that in our savings/income problem, the line crosses the y-axis
exactly at the origin of the graph (0, 0). In that case, when x = 0, y= 0 also,
and the y-intercept (@) of that particular line is 0.
LBM BAA AMA AONE NNN NTN mate ttn nieneiaisensitie

y-intercept The value of y at that point where the line crosses the y-axis
(i.e., where x = 0).
LMM OAL ON OLA OTIC ANN tUeaetatueNtiimmnsntmeeetee

The second constant in the equation of the line is designated b and is


called the slope of the line. It is defined as the change in y per (one) unit
change in x. Imagine two points on the line, P, with coordinates (x,, y,) and
P, with coordinates (x,, y,). The slope of the line is the ratio of the change
in y to the change in x, that is, the difference between the two y values
divided by the difference between the two x values.

aver ya
b
XD = Xi
Correlation and Regression Analysis j» 457

Figure 13.8

4 L

i y-intercept of |,
I

oa >
cea y-intercept of |,
ne y-intercept of |,

I; is parellel to the
y axis. There is no
y-intercept.

y-intercept of |, does exist, even though in order to find it


! one may need to graphically extend both the line and the y-axis
ce beyond the normal length of the graph or even off the page.
!
1
458 € STATISTICS FOR THE SOCIAL SCIENCES

Slope The change in y per unit change in x.

P,(&,, y;) and P,(x,, y,) may be any two points on the line; it does not
matter which ones. In our example, for instance, Jim’s two scores have
been graphed as a point on the line designated by the ordered pair (8, 4)
(see Figure 13.5). Let that be P,. Thus, x, = 8 and y, = 4. Jack’s scores
are designated by the ordered pair (16, 8). If we let this be P,, then x, = 16
and y, = 8. We can now find b.

b
Woe yr i ees
OG} 5) 16-8

By inspection of Figure 13.7, we see that every time we move a point up the
line far enough to change x by 1 unit, the corresponding change in y is 1/2
or .5 units.
To confirm our conclusion, we may designate another two points on
the line as P, and P, to see if we still get b equal to .5. Let us call Jill’s score
(20, 10) P, and Carol’s score (4, 2) P,.

b ye—-y1 2-10 -8
= eS -1
iD Xa 4 — 20 —16 —2

We can now indicate the equation of the line in our example. Remember
that the standard form of a linear equation is

y=a+bx

We know by visual inspection that a@ is 0, and we have calculated a b of .5;


thus,

y=0O+ 5x

To simplify the equation, when a = 0, we drop it altogether, Our equation


thus becomes

y= .5x

Earlier, we found the y values for Ann and Frank by visual inspection. We
now find them algebraically by plugging Ann’s and Frank’s x-values into the
formula and solving fory,
Correlation and Regression Analysis j» 459

For Ann:

y= .5(24) = 12
For Frank:

y= 5(10) =5

Since y was measured in thousands of dollars, we now know that Ann and
Frank saved $12,000 and $5,000, respectively.
A few more points are in order about b, the slope of the line. In our
sample problem, where b = .5, the slope is a positive number. However, in
other cases, the slope may be negative. The sign of the slope gives us an indi-
cation of the direction of the line and the nature of the relationship that the
line reflects. When the slope of a line is positive, the line will slope upward
from the lower left-hand side of the graph to the upper right-hand side (see
Figure 13.9). If the slope of the line is negative, the line will slope from the
upper left-hand side to the lower right-hand side (Figure 13.9). When the
slope of the line is positive, x and y are positively related. As x increases in
magnitude, so will y; as x decreases in magnitude, so [Link] our example
problem, the greater the amount of one’s schooling, the greater one’s
savings; conversely, the lower one’s schooling, the less one’s savings. If the
6 in the equation were negative, we would have two variables that were
inversely (negatively) related: As one variable increased in magnitude, the
other variable would decrease. Later, when we are working with the slope of
a regression equation, we will see that the slope has the same sign as 7

Inverse relationship — Relationship in which one variable increases in magnitude


as the other variable decreases.

To visualize the difference between positive and negative relationships,


imagine a room with an electric heater in it. The more power put into the
heater, the greater its output and the warmer the room becomes. Power
input and temperature are positively related: The more power, the more
heat. If, however, we replaced the heater with an air conditioner, the rela-
tionship between power and heat would be negative, or inverse: The more
power going into the unit, the /ess the heat in the room. Power input and
room temperature are still related to each other, but the nature of the rela-
tionship that was positive in the first case is negative in the second.
Therefore, we will need to be aware of the sign of the slope whenever we
undertake this kind of analysis.
460 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 13.9 Positive and Negative Slopes Correlation and Regression Analysis

The following lines all have positive slopes:


A A f

rar ©

The following lines all have negative slopes:

shed ew

LINEAR REGRESSION

Up until now, the problems discussed have been graphed as straight lines,
and a linear equation could easily be derived from the data. It would be rare
indeed, however, to find actual social data that could automatically be
graphed as a straight line. Several reasons account for this fact.
The first and most obvious reason the data may not graph as a straight
line is that the relationship may not be linear. Some other kind of curve is
best for describing such a relationship. Figure 13.10, for instance, presents
a somewhat refined picture of the relationship between respondents’ ages
and their levels of participation in athletic activities. For instance, ymight be
the number of athletic events held in the previous 12 months in which the
respondent actually played. The respondent’s age is treated as the indepen-
dent variable. There is a relationship between the two variables, but it is
curvilinear, not linear.

Curvilinear versus linear relationship A curved versus a straight-line relationship.


Correlation and Regression Analysis » 461

Figure 13.10 A Curvilinear Relationship

Participation
A

> Age

In a case like this, the equation best fitting the curve would not be linear
but would be more mathematically complicated. Finding such an equation is
beyond the scope of this text, but we can easily make the observation that
there is a relationship that is not linear by graphing the relationship and
inspecting the graph. It is an all-important step that is often left out of
research and yet could be of crucial significance to the social scientist. For our
purpose here, in the remaining procedures presented in this chapter, we will
assume that the underlying relationship we are studying is linear. If the data
do not graph exactly as a straight line, we will assume it is due to other factors.
Of those other factors, measurement error or procedural deviations can
be major reasons for deviation from linearity. For example, in the annual
income versus campaign contribution problem, we might have been relying
on the people to report their incomes accurately. If they estimate or even
falsify these figures, the data we have will only be inexact estimates of actual
income. How is the value of their contributions determined? Tax returns?
Verbal estimates?
In addition, other factors can keep the relationships in the social
sciences from being strongly linear. Variables other than the two variables
we are studying may be involved. For example, in the income/contribution
problem, people appointed to patronage positions may not all have been
the district attorney’s cronies. Some may have been appointed because they
possessed specific expertise not available elsewhere. Perhaps the one and
only arson specialist in the county was hired for her expertise alone and thus
felt less pressure to contribute to the boss. Or, perhaps a strong partisan
462 < STATISTICS FOR THE SOCIAL SCIENCES

contributed far more than expected out of ideological conviction alone.


Such factors would obviously prevent a perfect “real-world” linear relation-
ship between income and campaign contribution.
Some combination of error in measurement or other factors could lead
to a distortion of the overall data that, when graphed, does not result in an
exact straight line. The set of points may show a pattern that looks some-
what like a straight line but is not definitely linear. In such cases, we may
study the relationship by means of a technique known as linear regres-
sion, which finds a line that “fits” the scatter of data points in such a way as
to provide for any given value of x the best estimate of the corresponding
value of y. We are then able to predicty from x the same way that we did for
perfect linear relationships.

Linear regression Technique that finds a line that “fits” the scatter of data points in
such a way as to provide for any given value of x the best estimate of the
corresponding value of y.

Suppose we are studying the relationship between social alienation and


religiosity. We develop a scale, ranging from 0 to 100, to measure the extent
to which an individual may feel alienated from the nonpolitical structures of
society—friends, family, church, and other reference groups or their values.
The higher the score, the more the alienation felt. Religiosity is measured by
another scale, ranging from 0 (least religious) to 10. Our hypothesis is that
the higher one’s level of social alienation, the lower will be that person’s
intensity of religious belief.

x= y=

Social Alienation Religiosity


ae 10
30
oe
40
40 Oo
=I
CONS

The graph of this relationship, shown in Figure 13.11, is known as a scatter


diagram (or scattergram or scatter plot) because the data points are
scattered on the graph and not a perfect line.
nineties earssiaelssssutiatemmesin ennai chanscnacadaee ees
Scatter diagram (or scattergram or scatter plot) Diagram with data points that are
scattered on the graph and not a perfect line or other figure.
HOUSES RESIS ASSESS SERRE StS NET
Correlation and Regression Analysis » 463

Figure 13.11

10 ®

Sue j ® e
3 e

ap 6
g 4

Zz

0 — ar ie alae i= Sea >


0 5 10 15 20 25 30 3)5 40

Social Alienation

The scatter diagram indicates a relationship near enough to linearity


that we can imagine that, had all other factors (such as the influence of other
variables on the relationship) and all measurement errors been eliminated,
we would indeed have a straight line. We can use a technique called the
least squares method of linear regression to find the equation of the line
that best fits these points and that we will call the regression of y on «x.
This phrase informs us that the line we generate will be the line that enables
us to most accurately predicty from x.

Least squares method Technique that finds the equation of the line that best fits the
points of a diagram.

Regression of yon x Technique that informs us that the line we generate will be the
line that enables us to most accurately predict y from x.

To understand what the least squares method actually does, look at the
scatter diagram in Figure 13.12, where the regression line has already been
drawn in between the points of some hypothetical relationship. The actual
observations are P,, P,, P;, and P,. The line running between these points,
the regression line, has a set of points on it—P/, P5, Pj, Pj —which have the
same x values as their corresponding points P,, P,, P;, and P,. (The apostro-
phe is read “prime.” If P, is read “P-sub-one,” then P; is read “P-sub-one-
prime.”) We use a P with another numbered subscript because P, and Pare
related; in this instance, they share a common x-coordinate.
464 <@ STATISTICS FOR THE SOCIAL SCIENCES

Figure 13.12

Py (xq, ¥1)
P3 (x3, Y3)

¥3-Y3
Pi (x1, y4)

P5 (x2, ¥2)al ¥2- Ya


Ps (Xz, Y2)

What distinguishes P, from P/is that the y-value for the actual observa-
tion P, is different (in this case, larger) than the y-value of the correspond-
ing point P/that actually falls on the line. In fact, the shortesty distance from
P, to the'line 15 y; Sy. For P,, the shortest y distance to the line is y, ae
(It will be a negative number because the point is below the line; hence, y,
is larger than y,. But the absolute value of the distance is still |v, =¥,l .)
The vertical lines in Figure 13.12 show each of these distances, indicat-
ing the deviation of these points from the regression line. The regression
line is the line (found by means of calculus) such that the sum of the square
of all the y deviations from it is less than it would be for any other line that
one might construct between the points. Actually, we state this characteristic
with the following expression:

>> O,-9;) =a minimum


ND sale

iy —¥, is they distance from the point to the line. We square each such
distance, or deviation, to get rid of negative numbers. Then we add up all
the squared deviations. The number we get will be smaller than it would be
for any other line we might have constructed to pass between the data
points. We will never have to derive the calculus part of this problem our-
selves; it has been done for us already. But it is important to understand
what this regression of y on x actually does. We have generated a line that
Correlation and Regression Analysis » 465

minimizes they distances from the actual data points to the line. Thus, if we
predict y from x using this regression line, our prediction of y should be
closer to the true value of y than any other prediction we might have made
using any other line drawn between the data points.
To see this effect, examine Figures 13.13 and 13.14. In Figure 13.13, the
least squares regression for predicting yhas been generated from three initial
data points: P, (1, 3), P, (3, 0), and P, (6, 2). The distance from each point
to the line (y Bay is. thenrdeterinined to be 103,171, and468:tor PP.
and P,, respectively. Squaring each distance and adding up the squared dis-
tances yields a S* (y—y')? = 4.44. Any other line drawn through the three
points will yield a > @ —y)? greater than 4.44. The actual formula for
the least squares line is y = 2.1 — .13x. When a slope is negative, we write our
formula this way rather than y = 2.1 + (-.13) x. Figure 13.14 illustrates this by

Figure 13.13

y = 2.1 — 1.3x The least squres regression.


Sly - y’ = (1.03)? + (-1.71)? + (.68)
= (nO6 2.92 E46
= 4.44

There is no line that can be drawn through P,, P5,


and P; whose Y(y - y)? will be less than 4.44.

Se
466 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 13.14

y = 1—.1x is not the least squares regression.


Siva =) + Eh ey
= 4.41 + .49 + 2.56
=7.46

This, or any line other than the regression, has a


Y(y - y’? larger than that of the regression line’s 4.44.

showing another line. In this instance, }> (y -yP= 7.46, which of course
is larger than 4.44. The least squares line’s 5° (y ~y)= 4.44 is less than the
one for the line in Figure 13.14 or any other line through these points.
We must point out that the line used to predict y from x, the regression
of y on x, is not the line that would give us the best prediction of x from y.
This latter line, the regression of x on y, is obtained by minimizing the
sum of the squared x deviations instead of minimizing the y distances.
Figure 13.15 gives a visual portrayal of what we mean by minimizing the
x distances. The line we have drawn in that figure is the regression of x
ony, the line such that

y> (,-x,) =a minimum


Usually, we will not need to know the regression of x ony since we usu-
ally want to predict the dependent variable y from given knowledge of the
Correlation and Regression Analysis j» 467

Figure 13.15

independent variable x, and not vice versa. Still, the differences between the
regression of yonx and the regression of x on y will be conceptually impor-
tant to us. Only when we begin with a perfect linear relationship will the two
regression lines be identical to each other.
In any event, our eventual need is to find the line for the regression of
y on.x from our original data. To clarify this, we first make a small notational
addition to the general form for a linear equation.

y = a). pi b,x x

The subscripts on @,,, and 6,,, tell us that this line will be the least squares
regression of y on x. If we want to predict x from y, we would use the fol-
lowing notation instead:

ea Ay . by

And of course, @,,. does not equal a,,, and b,,. does not equal b,,, except by
coincidence.
Note that these are predictions of y, not the actual y values from which
we generated the regression formula. To stress this point, we place a
circumflex above the y (y read “y-hat” because the circumflex resembles a
hat) to indicate that this is an estimate of y. Recall that we did this with
sigma-hat, an estimate of sigma.
468 < STATISTICS FOR THE SOCIAL SCIENCES

sonnets hie mis SRS CTI EEE ECE,

Predictions of y The prediction or estimate of y from a specific value of x as


generated by the regression formula.

To find the regression of y on x, we will need to find a,, and b,,. Once
we know these two numbers, we will have the equation. The formulas that
will yield the slope and y-intercept will be presented a bit later. But first, we
need to return to the topic of the correlation coefficient.

The Correlation Coefficient

Before we plunge into calculating b,. and a,,, let us determine their
value to us in predicting y from x. We dothis by first finding the correlation
coefficient 7 and its square, the coefficient of determination.
First, let us understand the purpose of calculating ~ Pearson’s 7 is actu-
ally a measure of how close the point distribution comes to linearity. When
r is 0, it usually means that the points are randomly distributed throughout
the scatter diagram. Thus, the regression of y on x is a horizontal line with
slope equal to 0, and knowing the regression does not improve our ability
to predicty from x. In some instances, an r of 0 means the data are related
in a manner other than a linear one. In short, an 7 equal to 0 means the vari-
ables are not linearly related (see Figure 13.16), and so a linear regression is
of no predictive use.
By contrast, an r whose absolute value is 1.0 tells us that there is a perfect
linear relationship between the two variables, and the linear regression gives
us perfect predictive capability. The higher the value of 7; the closer the points
come to linearity and the better the predictive capability of the linear regres-
sion equation. Moreover, the sign on the correlation coefficient is identical to
the sign of the regression slope b,,.. Consequently, ifr is positive, the variables
are positively related; if 7 is negative, the variables are inversely related.
Now let us turn to the calculation of ~ Pearson’s r has a formidable
formula:

nee) ny xy — OFX)

J [n ox? — Oo x)*][n Vy? — OCy)?]


Although this appears difficult, with the aid of a calculator and the proce-
dure presented below, we can work it without too much difficulty, What we
must do is generate all of the components of the equation from the original
raw scores. Then we will separately calculate the three expressions
nyXY - O[O OLD) 2 = O94, andn yyy = oy)’. Then we will plug
the three numbers obtained back into the formula to find r.
Correlation and Regression Analysis »» 469

Figure 13.16

No linear relationship, r= 0:

random scatter strong nonlinear relationship

Perfect relationship, r= 1.0: r=-1.0

1 i}
| 1
| | e
\ I e i.
| e e | e
} e? | e
| A ® @ | e
e
‘one e
i} e
i}
i} i}
i} 1
i} 1
l l

r is positive ris negative

Moderate or partial relationship, r is between 0 and plus or minus 1.0

ris positive ris negative


470 STATISTICS FOR THE SOCIAL SCIENCES

Note that }°x and }\y are the sums of the original x and y scores,
respectively. The expression ()°x)’, read “summation of x, quantity
squared,” is simply }°x times itself. Likewise, ()°y)*, the “summation of y,
quantity squared,” is }°y times itself. To find }°x*, “summation of x
squared,” we square each value of x and then add up the values. (Note that
())x)’ is not the same as )°x*.) We will also need to square allyvalues and
add them up to get )°y*. To find }°>xy, we multiply each x score by its cor-
responding y score and add the products together. The number of cases
(people, places, or things) actually measured, 7, is found by counting how
many pairs of x andy scores we have.
Following are the steps elaborating this procedure. The original data
columns, labeled x = and y =, are shown first.

x= y=
Social Alienation Religiosity
25 10
30 9
op) 8
40 8
40 ri

Step 1. To the right of the y column, add three new columns: x*=, xy = )

and y7=.
Step 2. Square each value ofx and enter it in the x*= column.

Step 3. Multiply each x score times its respective y score and enter it in
the xy = column.

Step 4. Square each value of y and enter it in the y?= column.

Your data should now look like this:

— WS = x= =
PG) 10 625 250 100
30 D) 900 270 81
55 8 1225 280 64
40 8 1600 320 64
40 fe 1000 280 49

Step 5. Count up m, the number of cases (pairs of scores), and enter


itasm=_ (in this case, m = 5) to the left of the bottommost value
OLx.
Correlation and Regression Analysis j» 471

Step 6. Add each column up to obtain }°x, }oy, )ox*, }oxy, and doy”
Put these totals under a totals line below the bottommost score in each
column, and label each total as indicated below.
Step 7. Below the values for >x and }°y, square those numbers to
obtain ()
}x)’ and ()°y’) and enter these two numbers as labeled below.

You should now have

EX y i DCS XY = y =

25 10 625 250 100


30 ) 900 270 81
35 8 122) 280 64
40 8 1600 320 64
40 wi 1600 280 49

=> ) x= 170 Yiy=42 ox? =5950 Sixy=1400 )/y* = 358

Ouxy= G70) Ooy*= G2)


= 28,900 = 1764
Step 8. Using the figures calculated above, find n}°xy - ())x)Q9):
Step 9. Calculate n°x? — (Yo x)’.
Step 10. Calculate n)*y* — oy)’.

You should obtain the following:

ny xy — ee) (9) = 5(1400) — (170) (42)


= 7000 — 7140

= —140

ny x? — (Sox) = 5(5950) — 28,900


S29 S01 20,000)

O50)
2

ny Oo») = 5(358) — 1764


= 1790 — 1764

6
Step 11. Calculate Pearson’s r from the formula.
472 STATISTICS FOR THE SOCIAL SCIENCES

= ny ay) hemi
Jn ex? — xy? — oy?) V850) 26)
—140 —140
= 9417
/22.000 148.66

Step 12. Square r (before rounding) to obtain the coefficient of


determination.

? = (-.9417)*= .8867

Let us round off and summarize.

r=-.94 a large inverse relationship.


7 =.89 89 proportion (89%) of the variation in religiosity can be
accounted for by the variation in social alienation.

This relationship is large enough that few would dispute the usefulness of
moving on to find the regression formula, which we will do shortly. First,
however, a bit more discussion of 7” is in order.

The Coefficient of Determination

For now, perhaps it is best to say that the concept of the proportion of
variation explained by the independent variable is a way of specifying the
impact of the independent variable in explaining changes in the dependent
variable. There is a very specific explanation for the meaning of 7°, which will
be discussed in Chapter 14. Explained briefly here, variation is the sum of
all the squared distances of the points’ y values from the mean value of y.
The distances are squared to get rid of negative numbers. This sum of
squared deviations is then broken into two components: (a) the squared
distances from the mean y value to the regression line (explained variation)
and (b) the squared distances from the regression line to the point (unex-
plained variation). The closer the points come to the regression line, the
greater the amount of the total variation that can be attributed to the line
and the less left unexplained (see Figure 13.17).
In some ways, the coefficient of determination is a kind of police officer,
keeping us honest when we may feel the urge to inflate the importance of
our findings. Relatively few rs in research are .7, .8, or .9; rather, they tend
to be in the .2 to .5 range. To see what this means in terms of 77, look below
where rs and their 7’s are compared as 7 shrinks from 1 to 0.
Correlation and Regression Analysis p 473

Figure 13.17

®
Amount unexplained
by the regression line
Total
Variation
Amount explained —»
by the regression line
ST 4

If most points are this far from the line, r? will be moderate.

Amount unexplained
by the regression line

Total
Amount explained é Variation
by the regression line

r= ee
1.00 1.00
.90 81
80 64
70 £9)
.60 36
50 25
40 LG
30 09
.20 04
10 O1
0.0 0.00
474 @ STATISTICS FOR THE SOCIAL SCIENCES

By the time r drops to .7, ris only .49; only .49 proportion (49% — less
than one half) of the variation is explained, and .51 is left unexplained by the
independent variable. The proportion left unexplained (the reciprocal of
the coefficient of determination), 1 — 7”, is referred to as the coefficient of
alienation. By the time 7 = .40, only .16 of the variation is explained, and
the coefficient of alienation is .84; 84% is left unexplained.

Coefficient of alienation The proportion of variation left unexplained by the


independent variable.

At the same time, recall the earlier discussion of the many real-world
factors that mitigate linearity in the social sciences. It would be rare indeed
to see coefficients of determination approaching 1.0 (or even .9 or .8) in
social or behavioral research. Also, the concept of the proportion of vari-
ance explained assumes that no other variables are present that would have
an impact on the two variables in the relationship between y and x, but as
we saw in Chapter 6, that assumption could be unrealistic. Thus, perhaps it
would be better to say that the coefficient of determination is the propor-
tion of variation in the dependent variable that potentially could be
accounted for by the independent variable. For these reasons, while 77 is a
useful concept in statistics, it is less useful in the real world of data analysis
in determining what is or is not a “good” linear relationship. Thus, statistical
significance is more often used to demonstrate “good” relationships. This
will be illustrated in Chapter 14.

Finding the Regression Equation

Let us return to our social alienation versus religiosity example. With


r = —.94 and r’ explaining 89% of the variation, it is worth generating the
regression formula for predicting religiosity from alienation. If you found
the calculation of Pearson’s 7 to be tedious, here is joyous news: In calcu-
lating 7 we also obtained the essential components for finding bd...
According to the least squares principle, the following formula will yield
the regression slope:

ioe. n> xy — Oo» OO: y)


yx n> eee

In calculating 7 we already found that “xy — O ox) O"y) =-140 and that
n> )x* — ()°x)* = 850. Plugging these numbers into the formula, we obtain
Correlation and Regression Analysis j» 475

b, — ML - Ox) y) _ -140
= —.1647058
Ce = a50

The by just like x is negative, indicating an inverse relations


hip. Later
we will round off the slope to —.17, but for now, keep b,,. to
as many deci-
mals as your calculator allows. We need b,,. to find a, and even a slight
xe?

rounding of b,,. can have a large impact on Dyes


There are two formulas that may be used to find d,,. If the means forx
and y have already been calculated, we may use the following formula:

Ay ee b yx?
xe

If no means are available, we may use the following equation, which uses
information available from the initial calculation of Pearson’s r:

For our problem:

Viv -— Dx ox 42—(=.1647058)(170) 42 — (—27.999986)


Aayx = : = =
n 5 5
_ 42 + 27.999986 _ 69.999986 = 13.999997
5 5

rounding off: a, = 14.00. Thus, y= 14 + (-.17) x, which simplifies to our final


equation: .

y =14-.17x

Reminder: The negative sign before the slope indicates an inverse rela-
tionship. When the original relationship is positive, a + sign appears.
This is our best model for predicting the level of religiosity (v7) from
the social alienation score (x). Before we complete this procedure by
completing the scatter diagram, let us make use of our newly generated
equation. Remember that all of our work up to now has been to (a) estab-
lish the existence of a viable linear relationship by finding r and 7? and
(b) develop the regression formula for predicting y from x. Now that we
have the equation, let us put it to use.
476 @ STATISTICS FOR THE SOCIAL SCIENCES

Two things before we start. First, even though the original religiosity
scores ranged from 0 to 10, our estimates may be above 10 or below 0. We
are fitting actual data to a mathematical model. If the estimate is higher than
10, it simply means that were the index larger, this person would score even
higher than someone with a predicted score of exactly 10.
Second, the actual scores, based on only 10 items, produce only whole
numbers—10, 9, 8, and so on—but our estimate is a continuous variable
that can take on values between whole numbers: 9.8, 6.4, and so forth. If the
scale had more items, say 100 items worth 1/10 of a point each, then the
respondent could be expected to achieve the score estimated by the regres-
sion. If the scale only has 10 items, we could choose to round our estimate
to the nearest whole number between 0 and 10. In our example, however,
we will leave the estimates as the regression equation predicts them.
Now, to make an estimate using the regression equation, we simply
replace thex in the equation with the actual value specified for a person and
solve for y.
Suppose x = 0, no alienation:

J =14-.17x=14-.17(0) =14-0=14
We estimate 14 even though 10 is the highest possible score. Remember
that the definition of a,,, the value of y where the line crosses the y-axis, is
the y value where x = 0. Thus, (0, @,,.) is always a point on the regression
line. In this case, (0, 14) will be on the regression line—a useful thing to
remember when we add the regression line to the scatter diagram.
Now suppose x = 5:

J =14- 175) = 14=- 85 = 13.15


Thus, the ordered pair (5, 13.15) will also be a point on the regression line.
For x = 45:

3 =14-.17(45) = 14- 7.65 = 6,35


and so (45, 6.35) will also be on the regression line.
As an exercise, plug into the formula each of the five originalxvalues for
our social alienation versus religiosity problem to get the corresponding esti-
mates of y. Compare these estimates to the actual y score for each person.
We now want to add the regression line to the scatter diagram originally
presented in Figure 13.11. We do this in Figure 13.18, which has the same
data as Figure 13.11. Remember that any two points determine a straight
line; thus, once we know two points on the line, we can plot them
on the
graph, lay a ruler between them, and draw in the line. We already know
Correlation and Regression Analysis j» 477

three points on the line: (0, 14), (5, 13.15), and (45, 6.35). We plot the first
point (0, 14), (indicated with a dot in Figure 13.18) and then another point
within a ruler’s distance of the first. In this case, we will use (45, 6.55)
We then draw in the line, label our axes, add a title to the diagram, and
either along the line or at some uncluttered place on the graph, write the
equation, 7; and 7. We have now completed the entire correlation-regression
process. :
For further practice, let us return to the political corruption problem
presented earlier in this chapter, where we predicted monthly campaign
contributions from annual income. The following computations yield an
r = £81 and an 7 = .65. Although .65 is less than the 7? of the previously
worked problem, it is still quite satisfactory to justify finding the regression
formula. (Even though the findings for r and 7? might not be conclusive

Figure 13.18 Predicted Religiosity Score, by Social Alienation

oO @ (40, 8)

Religiosity
= (45, 6.35)

| eee! ime T >


=F ed

0 5 10 1 20 25 30 35 40 45
x = Social Alienation
478 <@ STATISTICS FOR THE SOCIAL SCIENCES

enough to convince the electorate {let alone a jury| that the district attorney
in question enforced an organized kickback scheme, .65 proportion of the
variation in contribution can be explained by income.)

x=
Annual y=
Income (in Mean
thousands Monthly
Employee of dollars) Contribution c= xy = y=
| 80 160 6400 12,800 25,600
70 2D 4900 6650 9025
3 52 oF 2704 5044 9409
4 45 85 2025 3825 V2n>

n=4 Meet, yaaa yox?= 16,029 iy’= 51,259


Soxy = 28,319 >
QO) =1247) Oly a7)
= 61,009 = 190,969

ny xy 2) (°y) = 4(28,319) — (247)(437)


= 113,276 — 107, 999
= 5337

n ce - (Sox) = 4(16,029) — 61,009


= 64,116 — 61,009
= 3107

ny y- (Sy) = 4(51,259) — 190,969


= 205,036 — 190,969
= 14,067

oe n> xy — OQ
X)O_Y) 7 5337
. J[n > x? — O¢x)2] [n y? — Of y)*] ~~G107)(14,067)
poll Boo
= +.807 = +.81
~ /.706,109 6611.06
P= C807) eGo) era
Correlation and Regression Analysis 479

pens n> xy—(x)Oly) = 5337


tfee Se ere ey

LLY — Pym IX _ 437 — (1.7177341)(247) _ 437 — 424.28032


Ayx =
n \" 4 2 4
= 12771968
= 3.17992
- 4

After rounding to two decimal places, we have our equation:

y= 3.184 172s

Note that unlike the religiosity vs. social alienation example worked
previously, which was inverse and had a minus sign in the regression equation,
here the relationship is positive, b is positive, and the + sign appears in the
regression equation. The scatter diagram is presented in Figure 13.19, where
the ordered pairs for each data point are indicated. (If 7 were very large, we
would probably not label the points at all.) Since a, = 3.18, we locate that
point on the y-axis as one known point on the regression line. To find a sec-
ond point, we select a value of x and solve fory . In this case, we chosex = 60.

y= ole + 157260) = 5.18 + 103.20


= 106.38

Thus, the point (60, 106.38) is plotted on the scatter diagram, and the
regression line is drawn in.

COMPUTER APPLICATIONS

SPSS

To find only the correlation coefficient for the social alienation/religiosity


problem, first key in the data as in Table 13.2.
Then, from the menu bar, click the following:

Analyze
Correlate
Bivariate
480 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 13.19

(60, 106.38)

e@ (70, 95)

Contribution
Campaign
Monthly
Via

0 10 20 30 40 50 60 70 80 90 $100

x = Annual Income (in thousands of dollars)

Table 13.2

VAROOOO1 VAROOOO2

1 25.00 10.00
2 40.00 9.00
2) 35.00 8.00
4 40.00 8.00
5 40.00 7.00
Correlation and Regression Analysis j 481

Table 13.3. Correlations: SPSS

VAROOOOL VAROOO002

VAROOOO1 Pearson Correlation 1 —.942*


Sig. (2-tailed) 017
ay. 5 5
VAROOO02 Pearson Correlation —.942°* 1
Sig. (2-tailed) (Oily ;
mM 5 5
a. Correlation is significant at the .05 level (2-tailed),

Move both VARO0001 and VARO0002 into the Variables box. This is really
all we need to do since we leave the correlation coefficients menu on
Pearson. (If we wanted to, we could also go to options and request means
and standard deviations for the two variables.) Click ok.
In Table 13.3, we see the output, a correlation matrix showing an
r of —.942. Under the correlation, it says sig. (2-tailed), which is .017. This
is a probability level associated with a test of statistical significance for % a
procedure to be discussed in the next chapter of this book. Note for now
that our probability is less than .05, indicating that the correlation coefficient
is probably not the result of sampling error.
We can also get the correlation coefficient when we get our regression,
although as you will see in a minute, care must be taken in its interpretation.
Much of what appears in the output will be strange to you until you have
read the next chapter. For now, ignore what is not specifically covered
below. (But after you have read Chapter 14, come back and reexamine these
results.)
Going back to the data list in Table 13.2, select the following:

Analyze
Regression
Linear

Now move VAROO002 to the Dependent Variable box and move


VAROO001 to the Independent Variable(s) box. Leave the other buttons
alone and click ok. The output is reproduced in Table 13.4.
Note that under Model Summary is our correlation of .94174. However, no
sign is given, even though we know from the above correlation run that the
relationship is inverse. If you use the regression procedure alone, it will give
you 7, but you will have to provide the sign yourself. In this case, examining the
482 <@ STATISTICS FOR THE SOCIAL SCIENCES

Table 13.4 Regression: SPSS

Variables Entered/Removed”

Model Variables Entered Variables Removed Method

1 VAROOO 1" Enter

a. All requested variables entered.


b. Dependent variable: VAROOO02.

Model Summary

Adjusted Std. Error of


Model R R Square R Square the Estimate

1 942° 887 849 4428

a. Predictors: (Constant), VAROOO01.

ANOVA?

Sum of Mean
Model Squares df Square F Sig.

1 Regression 4.612 1 4.612 25.520 a


Residual 588 BS) .196
Total 5.200 4

a. Predictors: (Constant), VAROOO1,

b. Dependent variable: VAROOO02.

Coefficients’

Unstandardized Standardized
Coefficients Coefficients

Model B Std. Error Beta t Sig.

1 (Constant) 14.000 AlASE T1950 001


VAROOO001 = llas, .034 —.942 —4.850 017

a. Dependent variable: VAROOO02.


Correlation and Regression Analysis » 483

data list is an easy way to spot that the relationship is inverse. However, most
data lists are not as user-friendly as this one, so be carefull! It is a good idea
to
run the correlation procedure as we did above, just to be on the safe side,
or get the sign from the slope. To the right of the R is R Square, which is our
coefficient of determination, .887. Of course, this is always positive. Skip
the
Adjusted R Square and the analysis of variance information for now; they will
be discussed in the next chapter.
Skip down to Coefficients and look under B. The number to the right
of our independent variable, VARO0001, is the slope, —.165, which confirms
that our r is also negative. To the right of (Constant) and directly above our
slope is the y-intercept of 14.000. We now have the information needed to
construct Our regression equation, y = 14 — .1@x, or if you round the slope
up as was done earlier, y = 14 — .17x. Now try doing the income versus con-
tribution problem on the computer.
If you wish, SPSS also provides a scatter diagram. From the original data
set, click on the following:

Graphs
Scatter

Define

Move VARO00002 to the y-axis box and VARO0001 to the x-axis box. Click ok.

Table 13.5 ~Correlations—SAS

The CORR Procedure


2 Variables: A B
Simple Statistics

Variable N Mean Std. Dev. Sum Minimum Maximum

A > 34.00000 6.51920 170.00000 25.00000 40.00000


B 5 8.00000 1.14018 42.00000 7.00000 10.00000

Pearson Correlation Coefficients, N =5


Prob > |r| under HO: Rho = 0
A B
A 1.00000 —0.94174
0.0167
B —0.94174 1.00000
0.0167
484 <4 STATISTICS FOR THE SOCIAL SCIENCES

SAS

As always,

Solutions
Analysis

Analyst

Use the upper-left icon, if necessary, to enter a new data set, and key in
your data in columnsA and B, just as was done earlier (see Table 13.2 for the
numbers). Then click

Statistics
Descriptive

Correlation

Or, at the top of the page, click the fourth icon from the right to auto-
matically go to the correlation program. Highlight A and click the correlate
button. Then do the same for B. NowA and B both appear in the center box.
Click the plots button and then click the box to the left of the word scatter-
plots. Click on ok. Back in the main Dialog box, click ok. Table 13.5 displays
the correlation results.
To view the scatterplot, there is a box on the left under the word reszi/ts.
Inside the box, see again results. Click on results and then on GPlot and
when the line Scatterplot of Ax B appears, double click on that. The diagram
will appear on the center of the screen.
At the bottom of the screen, click the analyst button to go back to the
original data. (Note: The correlations button would bring you back to your
correlation Output, and the Graph 7 button returns the scatterplot.) Once
you have pressed the analyst button, you can run the regression by clicking

Statistics
Regression

Simple

An alternate would be

Statistics
Regression
Linear
Correlation and Regression Analysis > 485

The latter can also be called up via the second icon on the right at
the
top of the screen. (The Simple regression procedure is fine for now. We’ll
revisit Linear in the next chapter.) Whichever way you go, highlight and click
B into the Dependent box and A into the Explanatory box. Click ok.
The
regression Output that appears is reproduced in Table 13.6.

Excel :

First, input the data from Table 13.2 underA and B. Then click on

Tools
Data Analysis

Correlation

Table 13.6 Regression: SAS

The REG Procedure


Model: MODEL 1
Dependent Variable: B

Number of Observations Read .)


Number of Observations Used 2)

Analysis of Variance

Sum of Mean
Source DF Squares Square F Value BP SIP

Model 1 4.61176 4.61176 Za) 0.0167


Error 3 0.58824 0.19608
Corrected Total 4 5.20000

Root MSE 0.44281 R-Square 0.8869


Dependent Mean 8.40000 Adj R-Sq 0.8492
Coeff Var Do WMS?

Parameter Estimates

Parameter Standard
Variable DF Estimate Error t-Value ee yi

Intercept 1 14,00000 1.17156 1195 0.0013


A all —0.16471 0.03396 —4.85 0.0167
486 << STATISTICS FOR THE SOCIAL SCIENCES

Click ok. Highlight the input range. $A$1:$B$5 appears. Click ok. The
correlations appear on Sheet 4 (see the left-hand side of the bottom of the
page). It is reproduced in Table 13.7.

Table 13.7. = Correlation: Excel

Column 1 Column 2

Column 1 1
Column 2 —0.941742 l

Click back to Sheet 1 to do the regression and click on

Tools
Data Analysis

Regression

Highlight row B and click to enter it in the Input Y Range: $B$1:$B$5.


Then click the cursor to the Input X Range box and highlight row A. In that
box, $A$1:$A$5 should appear. Click ok and the regression output appears
Omoheer s (see lable 13.3);
For a scatter diagram, click on the icon directly above the H in row H.
The Chart Wizard will open. Click xy scatter and then the finish button on
the bottom right. The diagram appears on the same screen as your data, but
if you print out the page, you will get the scatter diagram on a page by itself.

CORRELATION MEASURES FOR ANALYSIS OF VARIANCE


Once we have rejected the null hypothesis in an ANOVA problem, we
are concluding a difference in the population between at least two Of our
category means. In the nonexperimental context, that observation is tanta-
mount to concluding the existence of a relationship between the variables.
Once /7, is rejected, we may measure the amount of relationship in the data
and use that measure as an estimate of the relationship in the population.
There are two commonly used measures of association: the intraclass
correlation and the correlation ratio. The intraclass correlation coeffi-
cient, designated r, or sometimes 7, was once commonly used by
researchers. The simplest formula for calculating 7, is
Correlation and Regression Analysis j» 487

Table 13.8 Regression: Excel

SUMMARY OUTPUT

Regression Statistics

Multiple R 0.941742
R Square, 0.886878
Adjusted R 0.84917
Standard E 0.442807
Observation 5

ANOVA

af SS MS F Significance F

Regressor 1 4.611765 4.611765 ZO D2. 0.016732


Residual 3 0.588235 0.196078
Total 4 De

Lower Upper Lower Upper


Coef Std. Error — T Stat P-value 25% 95% 95.0% 95.0%

Intercept 14 PGS 58, 11.9499 0.00126 10.27158 17.72842 10.27158 17.72842


X Variable -0.164706 0.033962 -4.849742 0.016732 -0.272787 -0.050624 -0.272787 —0.056624

ae)
Se
n is the mean category size
F+(%-—1)

In our area versus pro-life score problem presented in Chapter 10,


we had two categories of four people each. Since both ms are the same,
n = 4.

os Nrural + “urban = 4+4 a 8 ll


th = ee
2; Zz ye

We review the source table to find that F = 10.80.

Source | SS ae | MS | F | D
Total \. 2800" * | | |
Between | 1800 | LF 1800 | 1080 | < .05
Within | 1000 | Gr #| 166.66 | |
488 << STATISTICS FOR THE SOCIAL SCIENCES

Thus,

al 1080-1 1080-1
9.80 — _
VY;
7 FS Gp = 0 ee Oe ee

There are several problems in interpreting 7, For one thing, a negative


r,does not necessarily indicate an inverse relationship in the data. Also,
when the category ms differ greatly, a more complex formula must be used
to find 7. Finally, although based on Pearson’s 7; 7, is more like a stand-in for
PROUT
More commonly found than 7, is a measure known as the correlation
ratio, designated by the capital letter E or by the Greek letter efa, n. We
always square the correlation ratio to make a generalization about propor-
tion of variance explained; thus, we really want to find E* or eta-squared (1°).

Intraclass correlation coefficient and the correlation ratio Two correlation


coefficients designed to measure association in an analysis of variance.

F? is readily calculated from information in the source table.

SST
For our pro-life problem, SS, and SS, are 1800 and 2800, respectively.

Thus, about .64 proportion (64%) of the total variation (total sum of
squares) can be explained by the categories of the independent variable.
The square of the correlation ratio is also often designated R?. (This is
similar to the k* known as the coefficient of multiple determination) to be
presented in Chapter 14. For now, simply interpret R? or E’ as if it were 7”.)
For instance, on the SAS Source Table printout (see Table 10.5), below the
source table, you will see R-SSQUARE and below it the number 0.244660. If
we calculate E* from the same source table, remembering that MODEL SS
means between SS in SAS usage,

SSeS
Correlation and Regression Analysis j» 489

So E’ is the same as R-SQUARE on the printout. In that problem, about


24.4% of the variation in course satisfaction can be explained by the
time of
the classes: day, night, or Saturday.

CONCLUSION
Correlation-regression techniques have become widely used throughout all
of the social sciences. Business and economics research also relies heavily on
this form of analysis. This is largely due to the predictive capability of regres-
sion. When r is large enough, we can rely on regression models for predict-
ing scores on such diverse dependent variables as electoral outcomes, voting
in legislative bodies, arms acquisitions, economic growth, unemployment,
health expenditures, academic performance, and crime rates.
Our predictions are enhanced when we turn our attention to an exten-
sion oflinear regression known as multiple regression, a topic to be encoun-
tered in Chapter 14. With multiple regression, we are predicting a score on
a dependent variable from several independent variables at once. As more
independent variables are added to the model, its predictive capability usu-
ally rises above the capability of a one-independent-variable model such as
we have been using here.

Chapter 13: Summary of Major Formulas

The Correlation Coefficient (Pearson’s 7)

a TO OFC)
Min ex? = (ox n hy? = Oy]
ea en oe FEe oe a, ,
The Least Squares Regression of y on x

ee Gyn. a DX

where

Dyx =
ney -— LYIOY ee a
> x = (ox)? n
490 << STATISTICS FOR THE SOCIAL SCIENCES

Correlation Measures for Analysis of Variance

The Intraclass Correlation Coefficient

fe.
fj, = ———— fis the mean category size
F+(n—1)

The Correlation Ratio é

SS

Use the following study to work Exercises 13.1-13.6 in this section.


A school psychologist has developed a series of inventories to measure student
interest in studying selected academic subjects. These inventories range from 0 to
100; the higher a student scores, the greater that student's interest in studying the
particular subject. The academic subjects being inventoried are health, physics,
athletics, algebra, literature, geometry, drama, and chemistry.

Exercise 13.1
For all 10th graders, the correlation between student interest in drama and student
interest in chemistry is —.90664. Below are the scores of the five 10th-pgiade repre-
sentatives elected to the student council.

Student DRAMA CHEM


Emile 14 93
Dimitri 8 93
Adam 7 100
Laura 0 85
Alexander 92 29
=

1. Calculate r for these five council representatives and compare it to r for all 10th
graders. Does the relationship hold for these five students? In case your calcu-
lator overloads:

[pe - (2) fev - (yy) = veszanrt6a0;


= 491547680
=22,170.8/5
Correlation and Regression Analysis > 491

2. For the council representatives, what proportion of variation in CHEM scores


can be explained by the DRAMA interest scores?

Exercise 13.2
1. Complete the regression analysis for the data in Exercise 13.1. Find b and a.
Assemble the regression formula.
2. Do a scatter diagram with the regression line included.
3. Compare the predicted CHEM scores with the actual scores of the five council
representatives.

Exercise 13.3
Here is a matrix of Pearson’s r correlation coefficients for the interest inventory
scores for all 11th graders in the same school. Review the coefficients.

HEALTH PHYSICS ATHLET ALGEBRA LITER GEOM DRAMA CHEM


HEALTH 1.00
PHYSICS 23 1.00
ALFILET A5 —.02 1.00
ALGEBRA 20 87 —.14 1.00
LITER = 12 —.83 Le —.83 1.00
GEOM 09 90 —.07 86 =e. 100)
DRAMA —.01 ei 01 —/2 90" = 92. 100
CHEM —.08 ae —.12 /0 = 69 09 — 92 1.00

1. Which two variables have relatively little correlation with most of the other vari-
ables? (Hint: These two variables are moderately correlated with each other.)
2. Assuming that all the others tap a verbal versus quantitative interest continuum,
which variables are positively associated with interest in quantitatively oriented
courses?
3. Which are positively associated with verbally oriented courses?
4. What are the highest positive and highest inverse correlations? (Ignore the r of
1.00 between a variable and itself.)

Exercise 13.4
Here are the regression formulas for predicting literature interest scores from several
of the other indicators.

LITER= 98.94 —.90 ALGEBRA r=—3


LITER = 87.34—.91 GEOM f= 05)
LATER 2117.22 — 1.27 CHEM r=—.89
LITER ==13.11 + .99 DRAMA r=-.90
492 STATISTICS FOR THE SOCIAL SCIENCES

1. Predict the literature score for 11th graders whose algebra interest is 0; 30;
80; 100.
2. One student has an actual LITER score of 5. His scores for each of the inde-
pendent variables are as follows:
ALGEBRA= 81
GEOM = 100
CHEM= 75
DRAMA = 23

Using the above regression formulas, predict his LITER score. In this case, what
index comes closest to predicting his actual score? Which is least close?
3. Suppose that Joan and David tied for the highest LITER score in the 11th grade
(100). Joan’s other scores are
ALGEBRA= 8
GEOM= 0
CHEM = 36
DRAMA = 100

Predict her LITER score from each of the regression formulas. Which is the closest
predictor and which is the least close?
4. David has the same ALGEBRA, GEOM, and CHEM scores as Joan, but his
DRAMA score is only 86. Predict his LITER score from the DRAMA score. Which
independent variables are the best and worst predictors, respectively, for David?

Exercise 13.5
When regression is performed by a computer program, so much information is
provided that one must hunt for what one needs. Following is a copy of an earlier
version of a SAS printout for predicting an 11th grader’s physics score from that
individual's geometry score. Before you examine the full printout and scatter
diagram, note that some of what you see, such as analysis of variance as applied
to regression, will be covered in the next chapter. What you need for now is found
in the lower portion of the printout, under the heading PARAMETER ESTIMATE (see
the partial printout below). The first number in that column, adjacent to INTERCEP,
is the y-intercept, a. Below that number, adjacent to the name of the independent
variable, GEOM, is the slope, b. (Ignore the two number ones under DF.)

PARAMETER
VARIABLE DF ESTIMATE
INTERCEP 1 41.122668 a
GEOM | 0.375980 b

See Figure 13.20 for the complete printout.


Correlation and Regression Analysis » 493

Figure 13.20

PLOT OF PHYS & GEOM LEGEND: A = 1 OBS, B = 2 OBS, ETC.

90
A] :
80 4 A
A AA
is a ALAS
: A AA ao
7O 4 A AK
A. A A A
A A aA
60 A

x :
PHYSae : BA 8 A
AA A
tod | A
Boa
B A
304”

ie
A
104
T T T cae neers t a Ea T T

0 10 50 60 70 80 90 100

In the scatter diagram, a pair of scores is indicated by the letter A instead of a dot.
If a pair of scores occurs twice, SAS prints a B. If three times, a C. The actual regres-
sion line is not printed here, though it is possible for the program to place a line of
letters approximately where the regression line goes. As an alternative, we could
add the line by hand by finding two points on the line using the regression formula,
exactly as we have been doing.
The modified scatter diagram appears in Figure 13.21.

From the printout, form the regression equation for predicting PHYSICS from
GEOM (use only the first two decimal places).
Use the formula to confirm that the two points used to find the regression
line in the scatter diagram are two points on the regression line: (0, 41.12) and
(100, 78.12).
Predict the PHYSICS score fer someone with a GEOM score of 55.

Exercise 13.6
See Figure 13.22 for the regression and scatter diagram printout for predicting the
LITER score from the GEOM score.
494 STATISTICS FOR THE SOCIAL SCIENCES
Correlation and Regression Analysis j» 495

1. Put together the regression formula after finding a and b on the printout. (Again,
use only the first two decimal places.)
2. Identify from the equation two points that could be used to find the regression
line on the scatter diagram. (Do not actually draw the line, except on a photo-
copy of the graph.)
3. You know from the matrix in Exercise 13.3 that the correlation between LITER
and GEOM is —.95. What proportion of the LITER variance can be explained by
GEOM: Should the regression be a good predictor?
4. Following are the actual LITER scores for three students. Predict their LITER
scores using the regression formula.

GEOM LITER LITER


Student Actual Predicted Actual
Julie 100 2 0
Jack 48 2 35
Jill 0 4 95

How close were your predictors?

Exercise 13.7
Calculate and interpret r, and F? for the ANOVAs in Exercises 10.2, 10.3, and 10.4.

Exercise 13.8
Using the computer, run the correlation and regression data for predicting Mean
Monthly Contribution From Annual Income, and compare your results to those
presented on pages 478 and 479.

NOTE
1. In most mathematics courses, they teach this formula as

= Tis 0

It is really the same information. Above, the slope is indicated by m instead


of b, which most statistical applications use. The b above is equivalent to our a.
The reordering of the equation makes no quantitative difference. Both mx + b
and b + mx yield the same number. When I was in college, I had three courses in a
row that all introduced linear equations at the same time: a math course (7x + b),
followed by a physics course (4+ bx), followed by an engineering course, which did
it another way, using something called vector notation. My grades weren't very
good that semester.
YW KEY CONCEPTS ¥

critical value Of 7/7 citical multiple correlation/R zero-order regression


spurious correlation coefficient of multiple slope
intervening variable determination/R? first-order (etc.) partial
causal modeling ordered triplet regression slope
partial correlation n-dimensional mnemonics
control variable hyperplane standardized partial
zero-order correlation multiple regression/ regression slope/beta
first-order partial multiple linear coefficient/beta weight/B
correlation regression adjusted R’/R’ adjusted
second-order (etc.) partial regression stepwise multiple
partial correlation slope regression
UMass IEEE LODE
ETE GE REE EEE LIEIIEENE EAE TEE OLENA RDI TE LEDS OETET EEESIME ELE SEDINE BEEE NELLIE ELIE LEE OEE TO LEE LENG NETS LEER OETA
CHAPTER

Additional Aspects
of Correlation and
Regression Analysis

Y PROLOGUE V¥

In the last chapter’s prologue, we talked about a relationship between


income and social status, with the latter variable dependent. But surely
other independent variables might also influence social status as well. What
might they be? Years of education? Prestige of the college one attended?
Social status of one’s parents? What if we could predict social status from all
of these variables and possibly others? Assuming that all our independent
variables are correlated with social status to begin with, wouldn't it be an
improvement in predictive accuracy to have more than one independent
variable at a time?
In this chapter, we extend what we covered in Chapter 13 beyond two
variables to three or more. Multivariate techniques, we call them. (Actually,
the bivariate techniques are a special case of the multivariate ones we will
encounter in this chapter, but let’s not quibble.) So we move from simple
correlation to multiple correlation and from linear regression to multiple
regression, and we explore a variety of related topics.
We only will scratch the surface of the multivariate universe in this
chapter. However, for you, the student, this will have to suffice for now. You
have reached the last chapter in this book. The course is almost over!
ANU LESLES—OD,

Pp 497
498 ¢ STATISTICS FOR THE SOCIAL SCIENCES

INTRODUCTION
In this, the final chapter of this book, we continue with the topic of
correlation-regression analysis and expand its techniques beyond two vari-
ables. In addition, a number of loose ends will be tied up. We will first seek
to tie together the two branches of statistics—descriptive and inferential—
showing how correlation coefficients and regression slopes may be tested
for statistical significance. We will pay particular attention to the analysis of
variance procedure as it is applied to a regression, as is commonly found
on a computer printout. We will also explain the origin of the critical values
of a correlation coefficient table.
The section on partial correlations and causal models presents a new
way of studying the impact of a third variable (the control variable) on the
relationship between two other variables. We first encountered this problem
in Chapter 6, when the use of partial tables was discussed. Now we will make
use of partial correlations to extend our analytical capabilities.
In the latter part of the chapter, we discuss multiple regression, the
extension of linear regression beyond a single independent variable. We
will make use of this and related concepts such as the coefficient of multi-
ple determination to develop and evaluate predictive models. These are
regression models aimed at enabling us to predict a score on a dependent
variable from the scores of several independent variables working
together.
We will see how such techniques may be used to verify theory and even
to help us develop further theory.

STATISTICAL SIGNIFICANCE FOR r AND b


When linear regression is performed by most library computer programs,
an ANOVA source table also appears and an F is generated. Here, analysis
of variance is being used to test a null hypothesis that b = 0. Because
where 6 = 0, 7 = 0 also, the null hypothesis is stating that in the popula-
tion, b=0 andr = 0. If 1, is true, any regression slope generated from the
sample and any Pearson’s r generated from the sample will differ from 0
only as the result of sampling error. If, however, the F.,,,,,.., in the analysis
of variance is statistically significant and we can reject H,, we conclude that
in the population, b # 0 and r # 0, and we use the sample regression/
correlation data to estimate their respective population parameters.
The conversion factors between the components of the ANOVA source
table and the original raw score data and 7? are summarized below.
Additional Aspects of Correlation and Regression Analysis j» 499

Source Sc df MS F
Total So = yp)"

Explained r2. iy =)? 1 ry —y)


; i)
(Between)
2 =
Unexplained (1-—77)-)iGW-—yp)* n—-2 hr s20) =) r(n — 2)
(Within) l-r

So from the regression, F is calculated as follows:

r2(n — 2)
F = ————._ _ df =1 and n—-2
1—r
To see what is happening, observe the deviation in they direction of some
point on the scatter diagram presented in Figure 14.1. In the y direction, the
distance from point P(x, y) to the mean value of y, y, is y —y. The distance
from any point to the mean value of y will be that point’s ycoordinate minus
y. Note that the distance fromy toy can be broken into two components. The
first of these is the distance, in the y direction, from the point to the regres-
sion line, which is y — y’, and the second is the distance from the regression
line to the mean value of y, which is y’ — y. Adding the two components back
together yields the original distance: (vy —y’) + (’-Y) =y —y" + —Y=y -V.
Of the total deviation y — y, the componenty’ — y is the part explained
by the regression line, and the remaining component, y —y”, is that part
unexplained by the regression line. If we square these deviations to elimi-
nate negative signs and add up all the squared deviations, we get Yo-y)y,
which is the same as the total sum of squares (in they direction). We do the
same for the two components ofy — y, and the following emerges:

WG apy =O =)) ta):


This is the same as saying

SSasa) = The portion of SS, The portion of SS,


explained by the ah unexplained by the
regression line regression line

and is analogous to

SSireayotal =SS Between I SSwyichin

except that “Between” means explained by the regression, and “Within” is


the error SS—the SS that the regression line does not account for.
500 STATISTICS FOR THE SOCIAL SCIENCES

Figure 14.1

When 6 =0 and r = 0, the regression line is actually the line parallel to


the x-axis where y =¥y. In effect, the ys are at y, each y’ — y = 0, and
>> 0” -y)* (the explained SS) drops to 0 (see Figure 14.2).
Sincey’—y = 0, }> (vy — 7)’ = 0, and there is no SS explained by the line.
Since SS, jainea IS the same as SS,, it stands to reason that

ue pn) eis
SSB eee0 tpt
afzg fp

and since

e -dises 0 ss
Mw My

iE a\()) when
r=0 and b=0

Let’s take a concrete example and do both the correlation/regression


part and the ANOVA. This time, though, we will use the definitional formu-
las instead of the computational formulas to calculate F. Assume the six data
points shown in Figure 14.3.

Figure 14.2

BiG 1
| r=Oand b=0
vam

ew Say
|

- y =y is the regression
(x,y) —> x=
Additional Aspects of Correlation and Regression Analysis j» 501

Figure 14.3

i=

4 bs « (1, 4) ° (2, 4) ° (3, 4)

3 ! re Nore: y =3
yee and every
2 e r r 13
| ie) (222) (82) :
1

0 a oe ee i! T re

1 2 3

y= = i = y= y=
1 2 | 1 2 4
2 Bus al 4 4 4
3 a | 9 6 4
1 7m 1 4 16
2 ee 4 8 16
iO 3 ee x 12 16
Yx=12 > py =18 Ea 28 oy 0 | jy= 00

ee 2 yas

n> xy- Do) (-y) SECC) ees zie = 16 ="


2
ny xt — ee) 2608) = a2? = 168 = 144-24

ny. v- Deal = 6(60) — (18)? = 360 — 324 = 36

oe i) 0 0
igen: no
0
= 35.39~°
_ aba — (0x) (Ly) = ° =0
i 0b ns
Yy-byx 18-2) _ 18-0 18
a
~ nN a 6 ar ret 6

20,0 =0,a=5
502 << STATISTICS FOR THE SOCIAL SCIENCES

The regression line is) =a +c =3 + O)x=3 + 0. Thus, ) =35. So here isa


case reflecting a null hypothesis: b= 0 or r= 0.
Now we will do the ANOVA using the definitional formulas. Remember:
SSwea = 2 -J)’, SSrrplained = DO =p). aud Se inexpuued = 0-9). And in
this problém;,y =3 and everyy =3:

x= y=

| Zz
2 2 J = 18/6 =3
3 2
1 1
Z 4
n=6 2 4
yas

VAV= O-PS =|) Pere GSS 1 Pera a


2 1 | 0 0 | = 1
=i (ae 0 er “4 1
= 17 | 0 Ow] 4 1
1 ion 0 Ge 1 1
1 1, | 0 oo 1 1
1 iNT 0 0 | 1 1
SG) =6 > 0’ -yy =0 Sway =6
= SSrtal a SSrexplained = SS unexplained

Confirming:

SSipeal Fa SStxplained 4 SSiinexplained

LO-JY=V 0" -F¥ + o-y?


6=0+6

6=6

To get MS. sainear WE divide SSexpiained PY Afgxpiainew Which is always 1. Think


of the categories as (a) explained by the regression line and (b) unexplained
by the regression line. Then (no. of categories — aS) hs Deh
Additional Aspects of Correlation and Regression Analysis j» 503

Figure 14.4

Regression

SS Explained ee 0 =)
MS explained =
af Explained 1

SSUnexplained
MS Unexplained = Ups
7 Unexplainec

Of inexplained = (Nora) — 1) — (no. of categories — 1)

=(6- 0) (2)ala

or, more simply,

Diver te aes dey 2

SO

SS Unexplained 6 6
= = 1S
MSy lained = SOS
aaa Af Unexplained 6-2 4

pe MS rxplained eA 0 2;
MSUnexplained Ls

By contrast, in Figure 14.4, there is a line where r #0, b #0, and F #0.
504 << STATISTICS FOR THE SOCIAL SCIENCES

nN
ea y —— | x? = XY —
S
1 ra.) 1 1
2 2 4 4
s ; | 9 2
1 Anil 1 3
Nn =6 2 4 | 4 8
3 5 | 9 15 ke
.O0
Oo
GS
Wi
il]
a

yix=12 >) y=i8 x" = 28 > xy = 40 pr |R

ny. xy- ee) (Sy) = 640) (12)8) = 240 — 216 = 24

ie ee bar = (e212) = ee — a
2
ny. ¥- (Sy) 16 (64) = (18 = 304 = 424 = 60

oe mye xy (ey) ee
le Sx? — (Sx)’] [»es (o9)"] 24)(60
J/(24)(60) ZA
1440

24 | ;

r* = (.63246)* = .40

_nyxy-(Xx) (Ly) _ 2 4 ale)


4
Ke =e)

ee) I—DOY ey
x 18 — Ct
(1)(12 ee 18og — | 12 EL
n 6 6 6

Thus, ¥= 1.0 + 1.0x or simply p=1+-.


Doing analysis of variance:

R I >
4 Il

=I = 18/6 =3

i =6

M4
Oy
Roe
RO
8 II
M= II- oS
ko
Go
Oo
BS
WW
x8
Additional Aspects of Correlation and Regression Analysis jb 505

y-V= (C-V/Y % | ¥ -7= Of =p? Sal yal = Oepre


By, 4 = im =i 1
=| 1 Onuel ET 1
0 0 | 1 ied e 1
0 0 =H Teal 1 1
1 Land 0 (seen 1 1
2 ee 1 Tegel 1 i
0 -FY = 10 Oi pid Lo-yV=
= SSrcat a SSrxplained = SS Unexplained

Confirming:

SSotat a Explained + SS Unexplained

10=4+6
10=10

SSExplained =f 4
MSkxplained =
Af explained 1

MS SS Unexplained 6 6 6 Ls
Winx (aii Clann — = — |
Ql inexolained Nie G2 4

MSexplained
= =| = (O60
MS Unexplained LD

Here, F # 0, although checking the table at df= 1, 4 we see that F< Fico
To confirm the relationship between F = 2.666 and r* = .40, we go to the
conversion formula.

Pin
Die
=2) =_ 406-2) =_ 404) _ 1.60 _ 6
= 4
Vai
EY 1 — .40 .60 .60

Before leaving this topic, we may now shed greater light on the
definition of 7° given in Chapter 13: 7° is the proportion of variation in the
dependent variable that can be explained by the independent variable.
If we replace the word variation with the more specific SS,,,,, and replace
“the independent variable” by “the regression line,” then

ry is the proportion of SS, »tal that can be explained by the regression


line.
506 << STATISTICS FOR THE SOCIAL SCIENCES

Or, simply,

ieee SS Explained = 20) =

SSToral “oO cit

SSy,,,, is the sum of the squared deviations in the y direction (vertical) of the
values of y of the points on the scattergram from the mean value of y.
SSryotainea 8 that proportion of SS;,,,,, accounted for by the regression line: the
sum of the squared deviations in they direction between the predicted ( ’)
the mean values y.

SIGNIFICANCE OF r

If we know Pearson’s r and wish to test its significance directly, we can


convert the 7 to an F using the previously developed formula F = [77(72 — 2)]/
(1 — 7’) and then compare the F to F.,,.3; at df = 1 and m — 2. However, it is
easier to use Table 14.1. The logic of this table makes use of an algebraic
transformation of the above formula to

2) F

7 yp oven

If we plug in forF the value ofF.,,,;.., at the .05 level, solve for 7*, and take the
square root of r°, we will find the lowest value of r that would be significant
at the .05 level at the designated degrees of freedom. Thus,

Fcritical
Feritical = ==
tae Po F critical

Table 14.1 gives critical values for either one-tailed or two-tailed tests.
(Since we always only use one tail of the sampling distribution of F, the
terminology really should be directional or nondirectional H,.)
In both cases, Ap: Population = 9. In the nondirectional (“two-tailed”)
instance, Hy: Wo wation # 9. If we could make a directionality assumption, we
would have as H, either Pcowie ) OC Tecmasuan, = Uy, wichever is. most
appropriate.
Let us assume a nondirectional alternative hypothesis for now. If 7 = 15,
then df = 1, 13, and at the .05 level, F critical —= 4,6/.

4.67 ey
Veritical = =. 246 = -514
15 = 2267, Bey noe.
Additional Aspects of Correlation and Regression Analysis 507

Table 14.1 — Critical Values of the Correlation Coefficient

Level of Significance for One-Tailed Test

O05 025 O1 .OO5


Level of Significance for Two-Tailed Test

df 10 : 05 2 01
1 988 997 9995 9999
2 900 950 980 990
3 805 878 934 959
4 729 ll 882 917
5 .669 754 833 874
6 622 707 789 834
7 582 Ce 750 798
8 549 Ge F16 765
9 521 602 685 735
10 497 576 658 708
1 476 553 634 684
12 458 532 612 ot
1B 44) 514 592 641
14 426 497 574 ae
15 412 482 558 606
16 400 468 542 590
17 389 456 528 S75
18 378 444 516 ai
19 369 433 503 549
20 360 423 492 537
21 352 413 482 526
We 344 404 472 515
23 337 396 462 505
34 330 388 453 496
25 eee 381 445 487
26 Si 374 437 479
oF Bil 367 430 471
28 3,06 361 423 463
29 301 355 416 456
30 296 349 409 449
35 O75 325 381 418
40 257 3,04 358 393
45 243 288 338 eT
a51 273 322 354
50 325
60 pail 250 295
70 195 Ey 274 3,03
80 183 One 256 283
awe 205 242 267
90 254
100 164 195 230
F. Yates, Statistical Tables for Biological,
SOURCE: Abridged from Table VIIof R.A. Fisher and
h (6th ed.), 1974. Reading , MA: Addison-Wesley, an imprint
Agricultural and Medical Researc
of Pearson Education.
508 STATISTICS FOR THE SOCIAL SCIENCES

At 13 df (n = 15), the lowest value of r that would be significant at the .05


level would be .514. Any sample r below that should not lead us to conclude
that 7 in the population differs from zero. By using the formula df=n — 2, we
get this information from Table 14.1 without needing to calculate FP. It shows
that at df= 13, at the .05 level, 7 sica = -514. We also see that at df= 13, the
Y ticats ACE 592 and .641 at the .02 and .01 levels, respectively.
If we had been able to make a directionality assumption in advance, we
could have used the one-tailed critical values of 7. At 13 df, we would now
only need an r of .441 to reject H,, as opposed to the .514 needed if no
directionality could be assumed.
Before we leave this topic, it is important to consider again the question
of when a correlation coefficient is large enough to have practical research
significance. In the previous chapter, this decision was based on the coeffi-
cient of determination and the proportion of variation in the dependent
variable explained by the independent variable. From a purely statistical
point of view, using the coefficient of determination is an appropriate guide
to establishing relevance. In most research we encounter, however, the
statistical significance of r is really what is used to “establish” research
significance, despite the fact that research significance and statistical signifi-
cance are quite different concepts.
The reason for this is that most of the correlation coefficients that we
actually generate in the social and behavioral sciences are relatively low. We
are much more likely to see coefficients in the range of .20 to .40 than in the
range of .70 to .90. Social and behavioral phenomena are complex, so we only
account for proportions of variation in snippets at a time. Remember, though,
that statistical significance is based in part on sample size, so the larger the 7,
the more statistically significant correlations are likely to be generated. Bear in
mind that some of these coefficients will not have as much research signifi-
cance as the researchers preparing such studies perhaps might wish.

PARTIAL CORRELATIONS AND CAUSAL MODELS

Up to this point in our discussion of correlation-regression analysis, we have


been dealing with relationships between two variables at a time. We now
turn Our attention to correlation techniques that involve more than two vari-
ables. Specifically, we are examining the effects of other variables on the
relationship between two given variables. In doing so, we revisit the three-
variable situations initially discussed in Chapter 6. Consider two variables,
designated x, and x,. (In this section, we will refer to each variable as x plus
a numbered subscript.) Suppose that the correlation between these two
variables, 7,, is .60. Now suppose that there is also a third variable, x,. When
we ask about the impact of x, on the relationship between x, and Xo, we are
Additional Aspects of Correlation and Regression Analysis » 509

asking to what extent the correlation 7,, is dependent on the presence of


variable x,. At one extreme, the presence of x, would have no impact on 7).
At the other extreme, the entire correlation between x, and, would be due
to the presence of x,; if x, were not present, there would be no correlation.
Recall that when a correlation between two variables is due to the presence
of one or more additional variables, we say that the relationship is indirect:
either spurious if the control variable is dependent, or providing interpre-
tation if the control variable is an intervening variable.

Indirect relationship — Relationship in which the correlation between two variables


is due to the presence of one or more additional variables.

Spurious Indirect relationship in which the control variable is dependent.

For example, suppose we are studying a group of volunteers who are


working for a charitable organization. We survey these volunteers and,
among other things, determine the following for each person.

x, =the mean hours per week that the subject is involved in volunteer
activities of a charitable nature.
x, = the dollars per year that the subject has donated to all charities.
x, = the mean annual income for that subject’s family.

We show the three variables and the Pearson’s correlations between


them in schematic form in Figure 14.5. By convention, the variable pre-
sumed to be the independent variable is placed on the left-hand side and
the dependent variable(s) on the right.
In a three-variable case, a quick way to determine if a correlation is
indirect is to see if it equals the product of the remaining two correlations.

Figure 14.5

70 x, Volunteer Work
Income x3 60

.80 x) Charity Dollar Contribution

M2 = .60

I'93 — 80
510 STATISTICS FOR THE SOCIAL SCIENCES

Figure 14.6 If r;, is indirect, then 7), = 743 °°? Here,'ry, °%3= (-70)C80) =
.56, very close to the actual r,, of .60. Since r,, is probably indi-
46 x rect, we indicate it with a dotted line on the schematic shown
x3 ! 60. —_—«in Figure 14.6.
80 me Note also that 7,‘ 7,, = .42 # .80 and te .48 # .70, SO
only 7,, is indirect. By saying 7,, is spurious, we imply that the
correlation is due to the presence of variable x,. The relation-
ship between time spent as a charity volunteer and money
Figure 14.7 spent as a charity contributor is due to the existence of the
respondent’s income. Note that if the original hypothesis had
a x been that higher incomes lead to greater involvement of all
% ' 69 _ kinds in charities, then the finding of r,, to be indirect would
80 ! confirm the hypothesis. If 7,, had not been indirect, we would
have to investigate the possibility that volunteer work moti-
vates one to contribute funds, or vice versa. Or perhaps both
activities motivate people to the other activity. Remember also
that we are only looking at three variables, and many other variables could
be able to influence charitable activities.
In all probability, however, income is an independent variable indepen-
dently influencing the two dependent variables, volunteer work time and
charity donations. Since in this case all correlations are positive, we can say
that the greater one’s income, the greater will be one’s time and money
spent in both activities. Since 7,, is indirect, volunteer work time is not
directly affecting charitable donations. Nor is it going the other way. We say
that the independent variable is causing the changes in the two dependent
variables. In Figure 14.7, we add arrowheads to the lines in our diagram to
indicate the probable direction of causation.
Now imagine a situation where the values of the correlations are the
same, but.x,, is now the respondent’s years of formal schooling. Assume also
that in time order, schooling came first, then the acquisition of income, and
then the time devoted to volunteer activity. The arrow now goes from x, to
x,, as seen in Figure 14.8.
Here, education is the independent variable, income is the intervening
variable, and volunteer work time is the dependent variable. Education

Figure 14.8

60 _- *1 Volunteer Work

Education 2 steel

X3 Income
Additional Aspects of Correlation and Regression Analysis > 511

“causes” volunteer work time, but indirectly, through the intervening


variable income. In effect, more education leads to higher income and
higher income frees up more time for volunteer activities. Income provides
interpretation of the relationship between education and volunteer work.

Intervening variable A [Link] must be present for the independent variable


to bring about change in the dependent variable.

When we eliminate indirect correlations and at the same time make


assumptions about causation, as we have just done, we are undertaking what
is known as causal modeling. The direction of causation (the arrows) are
based on logical assumptions, primarily time ordering. Despite the fact that
reciprocal causation might also be logical, we usually ignore that possibility
and place only one arrow on a given line. (That is, we ignore the prospect
that while more education leads to more income, more income may also
enable a person to acquire more education.) After setting up the model, we
eliminate “paths” by determining which correlations may be indirect.

Causal modeling Procedure in which we eliminate indirect correlations and at


the same time make assumptions about causation.

The Role of the Partial Correlation Coefficient

In the previous examples, we determined that 7,, was indirect by com-


paring it to the products of the other two correlations. Another approach
involves the calculation of partial correlations. In simple terms, the partial
correlation between two variables is the amount of relationship not attrib-
utable to other variables in the system. In our original problem, the partial
correlation between x, and x, is the amount of correlation not due to x,. It
is the amount of direct correlation shared solely by x, and .x,. When we find
the partial correlation between x, and x,, we say that we are controlling for
the effects of x, on that relationship, and we call x, the control variable.
We indicate the partial correlation as r,,,,with the number of the variable to
the right of the period being the control variable. Similarly, the partial cor-
relation between x, and x, controlling for x, would be indicated 7,,,. Thus,
we have new notation and terminology.

Partial correlation The amount of relationship not attributable to other variables


in the system.
Sy < STATISTICS FOR THE SOCIAL SCIENCES

emanceter SARITESSCECNORETHOEESTRU n MOHA RE9PARDIHIP LAESSEEHEET INL LE ES OOTE ELLLISETCEEISS SEMTSEET TETED,

Control variable The variable whose effects on the relationship between the other
variables is eliminated by finding a partial correlation.

The simple Pearson’s r between x, and x,. Because we are not controlling
for any other variables, this is also called a zero-order correlation.
The partial correlation between x, and x, controlling for the effects
of x,. Since we are controlling for only one variable, we refer to this
as a first-order partial correlation.
, The partial correlation between x, and x, controlling for the combined
effects of two other variables, x, and x, (a new variable in our system).
Since we control for two variables here, we call it a second-order
partial correlation.

Zero-order correlation Correlation that does not control for any other variables.

First-order partial correlation Correlation that controls for only one variable.

Second-order partial correlation Correlation that controls for two variables.

We can have as many control variables as we want, the order being the
number of control variables being used.
Calculating partial correlations can be complex.
Third-order partials are calculated from second-order
Figure 14.9
partials, second-order partials are calculated from first-
order partials, and first-order partials are calculated from
Zero-order (simple) correlation:
the zero-order correlations. Generally, we use computer
70 xy programs for this technique.
.60 However, in the case of our three-variable problem,
60 X the calculations are not hard. We need to find first-order
partials from the original zero-order correlations shown in
Figure 14.9. The formula for 7,, 12.3), is

(ip —WAleyeee a .60 — (.70)(.80) 7 .60 — .56


Oa ee
a 7, Vi 7, Vi=.70V1—
80 J1—
V1 = 64

04 4 ~~ 04
= .09335
S5l/36 CUIADCSY 42826

Rounding, we get 7,,, = .09.


Additional Aspects of Correlation and Regression Analysis 513

Whenever a partial correlation is close to zero, we say that the original


relationship was indirect. We went from .60 to .09. (A zero in the tenths
place + anything in the hundredths place can be thought of as nearly
0.)
Thus, .09 being close to 0 suggests a spurious correlation or other indirect
relationship.
We use similar formulas to get the other two partials.
~

TAS = 12 23 ./0— (60) G80) f


Ne ee = .46
ee 5 Wee V1 — .607/1 — .802

VIE, AGS .80 — (.60)(.70)


7231 = = = Gy)
jl Z Ae PE O0-A VERT 02

Summarizing, the first-order partials are presented in Figure 14.10


Figure 14.10. Only r,, ,approaches 0, so only one indirect
relationship is discovered here. If 7,,, had not been low, First Order Partials:
we would have concluded that there were no indirect
46 xX}
relationships. If that had been the case, we would keep i}

a solid line between x, and x, and put an arrowhead X3


1
09

on that line. In the example where x, was charity 67 X)


dollar contributions, this decision would be a difficult
one: Do contributions lead to volunteer work, or vice
versa? (See Figure 14.11.)
In the case where x, was education, the decision would have been
easier, assuming education occurs first in time (see Figure 14.12). Here,
education has a direct impact on volunteer work as well as an indirect
impact via income, the intervening variable.
Sometimes, when we generate partial correlations, we encounter a situ-
ation where the zero-order coefficient is very low, but the partial correlation

Figure 14.11

Either or

eo Volunteer x, Volunteer
1 | Work
Work Income x; ae
Income x; ae
x
Charity x
Charity
ae
Contributions Contributions
514. STATISTICS FOR THE SOCIAL SCIENCES

Figure 14.12

x; Volunteer Work

Education x {
X3 Income

is not. Here is an example. We assign attitude scores to survey respondents


as follows:

Pro-choice. The more pro-choice the respondent, the higher the


score.

Pro—death penalty. The higher the score, the greater the respon-
dent’s support for the death penalty.

X;: Religiosity. The higher the score, the more religious the respondent.

In Figure 14.13, we summarize our results.


Although initially, it appeared that there was no correlation between
pro-choice and pro-—death penalty attitudes, when we control for reli-
giosity, a large positive correlation emerges. Such situations could result

Figure 14.13

Zero-Order (Simple) Correlation

-.70 X1 Pro-Choice
Religiosity x3 105

-63 X> Pro-Death Penalty

First-Order Partials

~.94 X1 Pro-Choice

Religiosity x, 89

x2 Pro-Death Penalty
Additional Aspects of Correlation and Regression Analysis ® 515

from offsetting consequences of religiosity. If, for instance, among the


very religious, the relationship between x, and x, was the opposite
of the relationship for the less religious, the two effects would offset
one
another. Suppose for the very religious, respondents are either pro-choi
ce
and opposed to the death penalty or pro-life and support the death
penalty. For the nonreligious, respondents are either pro-choice and
Support the death penalty*or pro-life and oppose the death penalty.
Without knowing religiosity, this strange relationship would not be appar-
ent. It would be suppressed.
We can illustrate this using tables. Suppose 50 people have high
religiosity and 50 have low religiosity.

For the Religious


Death Penalty Pro-Choice Pro-Life
Support 0 LS | 25
Oppose 25 0 | 25) 5 O10)
S eS ue ae
For the Nonreligious
Death Penalty Pro-Choice Pro-Life
Support 25 0 | ZS
Oppose ue 25 | 25 OQ=+1.00
25 25 | 50.

But when we combine both groups and do not differentiate by religiousity,


the following table results.

Death Penalty Pro-Choice Pro-Life


Support 25 25 | 50
Oppose 2 25 | 50 pO 0
50 50 | 100.

We can also use causal modeling with some of the other coefficients
in this chapter, such as partial regression slopes and standardized partial
regression slopes (known also as path coefficients). These techniques
were more commonly applied in the 1960s, seemed to decline in popu-
larity until a few years ago, but appear to be staging a comeback in recent
literature.
516 STATISTICS FOR THE SOCIAL SCIENCES

MULTIPLE CORRELATION AND THE


COEFFICIENT OF MULTIPLE DETERMINATION

The concept of the multiple correlation, k, and the coefficient of multiple


determination, R?, is an extension of Pearson’s r and 7~ to more than two vari-
ables. In fact, some systems of notation and many computer programs use
R and R’ even for the simple bivariate r and 7°. The multiple correlation
coefficient measures the correlation between a dependent variable and the
combined effect of other designated variables in the system. The coefficient
of multiple determination measures the proportion of variation in the
dependent variable accounted for by those other variables. In the three-
variable example where x, is the dependent variable, R,,, tells us that x, is
dependent and that there are two independent variables, x, and x,. If you
see R,,,,, you would know that x, is dependent and that there are three
independent variables, x,, x,, and x,. If you see R,,, you know that x, is
dependent and x, and x, are the two independent variables. Similarly, R3 ,, is
the proportion of variation in x, accounted for by x, and x, acting together.

Multiple correlation Coefficient that measures the correlation between a dependent


variable and the combined effect of other designated variables in the system.

Coefficient of multiple determination Coefficient that measures the proportion of


variation in the dependent variable accounted for by those other variables.

Actually, the multiple correlation is simply a mathematical extension


of Pearson’s r. In the two-variable case, there were two perpendicular axes,
and the data points were based on coordinates referenced to the two axes.
The regression was the best-fitting line running through the data that
could be used to predict the dependent variable. When a third variable is
added, a third perpendicular axis for the new variable extends out from the
origin. Two- and three-dimensional coordinate systems are illustrated in
Figure 14.14. (The negative-valued axes are deleted for visual clarity in this
figure.) As in a hologram or a 3-D film, the points that were once on the flat
two-dimensional x, + x, coordinate system are now extended out toward
the viewer (or back behind the “screen” where x, coordinates have negative
values). Each point in space is referenced to the three axes by an ordered
plete, Xs, 0)

seaman
iaiaemeeanmemenmummmnnnmmmmmarennenuncneerant
ue ER TTT

Ordered triplet Each point in space that is referenced to the three axes.
Additional Aspects of Correlation and Regression Analysis j» 517

Figure 14.14
a.. ee
. eee

AD
A

Two Dimensional
Each point is indicated if.
by an ordered pair: op pies ce ad ; en)
(x1, X>)
!
a sae foseeen 1, 3,2)
| |
l |
i aioe hile!

= =
! |

?
= > x
1 7. 3

49)

Three Dimensional
Each point is indicated
by an ordered triplet: (3, 2, -2)
(x1, XO, X3)

Although this three-dimensional scatter of points in space, like its


two-dimensional cousin, could have no pattern (a “blob” of points) or could
be curvilinear, we are here still looking for linearity. In three dimensions,
a linear relationship becomes a plane. Think of the plane as a flat piece
of window glass with the data points scattered above and below it. The
multiple correlation coefficient is a measure of how close the data points
in space come to delineating a plane.
518 @ STATISTICS FOR THE SOCIAL SCIENCES

Figure 14.15 shows, in a simplified fashion, the relationship of the


points to the plane. When points are scattered far and unevenly around
the plane, R is low. As the scatter becomes closer and more even about
the plane, R grows larger. When the points all fall exactly on the plane,
R equals 1. Note how this parallels what takes place in two-variable linear
regression where, as the points become more linear, simple 7 increases,
approaching +1.0 as a limit.

Figure 14.15

Relationship of the Points to the Plane


as R Increases in the Magnitude

Data points are far


from the plane:
R is small

Data points are closer


to the plane:
R is larger

Data points are


actually on the plane:
R= 1.00
Additional Aspects of Correlation and Regression Analysis 519

We are not limited here to three variables; however, beyond three


dimensions, we lose the ability to graph the structure. Nevertheless, we can
still describe the structure mathematically and call it an n-dimensional
hyperplane. No matter how many dimensions, we can still find R and R’.
So we have the freedom to add as many new variables as we wish.
~

n-dimensional hyperplane A structure with more than three dimensions that can be
described mathematically.

Unlike the bivariate Pearson’s r, the multiple correlation is always a


positive number. It does not attempt to indicate whether the independent
variables are positively or inversely related to the dependent variable, simply
because when there is more than one independent variable, some may be
positively related, whereas at the same time, others may be inversely related.
Since R has no sign, its real value comes once it is squared and is indicating
the proportion of variation being explained.
As with partial correlations, the calculation of R and R* can become
complex as more independent variables are added to the soup. In the case
of a three-variable problem, the formula with x, dependent and x, and x,
independent could be either
Ze Ss 2 tes
Rio3 =1 12 T+ 32 A -Ti2)

Or

goal Bard >


hee
Rix = + M23 A-14

Verbally, using proportion of variation terminology, we take the proportion of


variation explained by one independent variable and add to it the product of
the additional proportion of variation explained by the second independent
variable times the proportion of variation left unexplained by the first inde-
n.)
pendent variable. (This is not an easy definition to grasp at first presentatio

Figure 14.16

_ » X; Volunteer Work

Education x << e |
X3 Income
520 << STATISTICS FOR THE SOCIAL SCIENCES

In the case of the earlier example where volunteer work time is being
explained by education and income together (Figure 14.16):

We Were Given We Found


11,= -60 r1,,= 09
Pig mind YD gg .46
r,,= .80 1p, = .07

Using the first of the two multiple Rk’ formulas, we get

ae = ae + eel a Ti)
(.60)? + (.46)*[1 — (.60)7]
= 36 + .2116(1 — .36)
36 + .1354
II 4954

About 492% of the variation in volunteer work time can be explained by the
combined effects of education and income. (The other R-,, formula yields
.4944, The difference is the result of rounding back the partial correlations
to two decimal places.)
Suppose you wanted a variable other than x, to be dependent. You can
make use of the same formulas by temporarily designating the dependent
variable as x, and the other two variables, in no particular order, as x, and Lee
Before we conclude this topic, let us address an often-asked question.
Why not add up the simple 7’s to get the multiple R’? In other words, could
we not say

It turns out that this formula is correct only if.x, and x, are uncorrelated; it
only works if 7,, = 0. Since this is generally not the case, R, ,, is almost always
less thiatie ne.

MULTIPLE REGRESSION
Multiple regression, a shortened term for its more formal title of multi-
ple linear regression, is the technique of developing predictive equations
when there is more than one independent variable present. Just as we can
find R and R* where we add variables, we also can develop a predictive equa-
tion to predict a score on the dependent variable from the combined effect
Additional Aspects of Correlation and Regression Analysis j» 521

of the independent variables. Recall that the equation for a straight line
followed the format y =a + bx, where a is the y-intercept and b is the slope.
If we indicate all of our variables as x + a subscript as we have been doing in
this chapter, then for x, dependent,

= a, + OX,
~

The new subscripts for a and 6 indicate that a, is the value of x, where the
regression line crosses the x,-axis, and b, is the slope going with variable x,.

Muitiple regression or multiple linear regression Technique of developing


predictive equations when there is more than one independent variable present.

When we add another independent variable, x,, our equation becomes


the equation for a three-dimensional plane:

Sen ,
Xa te, eX,

The prime symbols now above a, and b, are to remind us that these
numbers are not necessarily the same as the ones in the first equation. They
change each time a variable is added. If we bring in another variable, x,, the
equation for the four-dimensional hyperplane would have

eae DX Den
gs uM Vee Pins Le

As before, the as and bs change from what they were prior to the addition
aPx pe
Once we are at three dimensions and are dealing with a plane, the a
value becomes the intercept of the dependent variable, that is, the value of
the dependent variable when the two. independent variables’ scores are
each zero, which is the point where the plane crosses the dependent vari-
able’s axis. This is illustrated in Figure 14.17. If x, is dependent, we move
along the x,-axis to where the plane crosses it. Then, a/ is the distance of
that point from the origin. (If x, were dependent, we would move up the
x,-axis and intercept the plane at @,.)
The slopes in the multiple regression, the bs, are referred to as partial
regression slopes. We use the same terminology as we did with partial
er
correlations. The simple two-variable regression slope is a zero-ord
regression slope. When we add a third variable, the slopes become
the
first-order partial regression slopes. When we add a fourth variable,
ion slopes, and [Link],
slopes become second-order partial regress
522 << STATISTICS FOR THE SOCIAL SCIENCES

Figure 14.17 The Intercepts for x,, x,, and x,

XQ

Partial regression slope Slope in a multiple regression.

Zero-order regression slope A simple two-variable regression slope.

First-order partial regression slope A regression slope between a dependent and an


independent variable where the third variable is the control variable.

The order’s number tells us the number of other variables whose


influence is being held constant. To understand what we mean by “holding
constant,” examine Figure 14.18. If we wanted a simple regression slope, as
before with linear regression, we would take the original data points (not
shown here) and use them all (ignoring x,) to find the slope of the line
minimizing the squares of x, distances. With the partial regression slopes, we
are not using all the data points at the same time. Instead, what we do has
the effect of taking some specific value of x, (it does not matter which
Additional Aspects of Correlation and Regression Analysis j» 523

Figure 14.18 A Partial Regression Slope

XQ

X3
go
Each one of these cuts in the plane, where x, = 1, x, = 2, and x, = 3, respectively, is parallel
to the others, thus having the same slope. This is the partial regression slope of x, on x,,
controlling for (holding constant) the effects of x,. It differs from the simple (zero-order) b,,,
because that value is based on al/ points on the plane, disregarding their x, values.

value) and using only the data points with that particular value. Notice in
Figure 14.18 that if we take just the points on the plane where x, = 1 and
reflect them back on the surface where the x,- and x,-axes cross, those
points will compose a line. If we use only the points where x, = 2 on the
plane and reflect those points back, we will get a second line. Doing the
same where x, = 3, we get a third line (see Figure 14.19). All three lines,
when reflected back on the surface where the x,- and x,-axes cross, will be
parallel, thus sharing a common slope, which is the partial regression
slope. But that slope will not necessarily be the same as the slope estab-
lished when all the original data points were used to find the simple 0,,.
(Another explanation of b follows later in this chapter.)
All of this notation and graphical representation takes some getting
used to. However, even if we have difficulty visualizing all of these as and bs
and what they do, we can still apply multiple regression techniques. We can
trust that the equation we obtain will be the best equation, using the least
524 STATISTICS FOR THE SOCIAL SCIENCES

Figure 14.19 Reflections of the Cuts in the Plane (Figure 14.18) on the Surface
Where x,- and x,-Axes Cross

—— eo

<— Points where x3 = 1


<— Points where x3 = 2

nab Points where x3 = 3

All lines are parallel and


have the same slope.
| This slope is the partial
regression slope.

squares criterion, for predicting the value of a dependent variable from the
combined effect of several other variables. What makes our job easy is the
computer. Because the calculations for multiple regression intercepts and
slopes are tedious and complicated, they rarely will be done even with a cal-
culator. Let us, therefore, assume that we will have at our disposal a software
package that will calculate all items necessary to build the regression equa-
tion. Using a computer also enables us to disregard a lot of the subscripts
and other notation that have been used up to this point.

An Example From Judicial Behavior

Suppose some organization wishes to influence the process whereby U.S.


Supreme Court justices are selected. The organization has rated each judge
currently serving on the Federal Appeals bench or on a state supreme court
and has assigned each judge a score designed to measure support for indi-
vidual civil liberties. The scores range from 5 to 30. The higher the score, the
greater the judge’s judicial support for the “liberal” decisions of the Warren
Additional Aspects of Correlation and Regression Analysis 525

Court. The lower the score, the more that judge took a “strict construc
tionist”
stance in Opposition to the decisions of the Warren Court. Thus,
the more
“conservative” the judge, the lower would be his or her overall
Score;
You are a scholar in the field of judicial behavior and have been study-
ing many of these same judges. From previous court decisions, you
have
been able to code their positions along several areas of recent legal con-
tention. For each issue area, you have scored each judge on a scale ranging
from 1 (most strict constructionist) to 5 (most civil libertarian/liberal). You
have coded the overall score and each issue area score for each judge and
entered the data as computer input. The variables are indicated with the
short names listed below.

Variable
Dependent: Short Name
Judicial Liberalism JLIB
Independent:
Abortion/Pro-Choice ABOR
Capital Punishment/Opposition CAPP
Censorship/Freedom of Expression CENS
Consumer Protection CONS
“Right to Die” Legislation RDIE

In most software packages, the short names (mnemonics) replace algebraic


letters such as x, or x, in identifying the variables.
MEET

Mnemonics Short names that replace algebraic letters in identifying the variables.
NN

As the researcher collecting these data, you would not have been oper-
ating in a vacuum. Your expertise in judicial behavior would have led you to
select independent variables that you would expect to explain overall judi-
cial liberalism. You would generally not code those issue areas that, in your
experience, do not correlate with overall liberalism. If you had not previ-
ously noted a liberal/conservative pattern to issues such as import-export
regulation, zoning authority, or income tax law, you would exclude them
from your study. On the other hand, if there were no experiential or theo-
retical basis for excluding the latter issue areas, you might include them as
well. If they provide little predictive capability, the regression and its accom-
panying statistics will demonstrate that fact. Finally, if there is little theory
but many independent variables, you might make use of a procedure known
as stepwise regression, which will be presented shortly.
526 € STATISTICS FOR THE SOCIAL SCIENCES

From the regression program, you determine what later will be shown
to be the “best” formula for predicting JLIB. It predicts JLIB from the first
three independent variables on the list—ABOR, CAPP, and CENS:

JLIB = —.82 + 2.46 ABOR + 1.81 CAPP + 1.59 CENS

The intercept (sometimes referred to as the constant) is —.82. The partial


regression slopes are 2.46 for ABOR, 1.81 for CAPP and 1.59 for CENS. To
predict a judge’s overall score from the issue area scores, take the latter
scores and substitute the scores for the short names in the equation, just as
we did with the x values in linear regression. (Note that these coefficients
are not based on an actual study but are from simulated data.)
Let us begin by examining four hypothetical cases. A judge who scored
most conservative on each of the issue areas (a score of 1) would have the
following predicted overall score:

JLIB = —.82 + 2.46 ABOR + 1.81 CAPP + 1.59 CENS


= —.82 + 2.46(1) + 1.81(1) + 1.59(1)
==—82 + 246+ 1.81 + 1:59
= —,82 + 5.86
= 5.04

A judge who scored most liberal (5) on each issue area would have this
predicted overall score:

JLIB = —.82 + 2.46(5) + 1.81(5) + 1.59(5)


= —.82 + 12.3 +9.05 + 7.95
= S48

(Recall that the original scale ranges from a low of 5 to a high of 30.)
The next case is a relatively liberal judge. His scores are 3 on ABOR, 4 on
CAPP, and 5 on CENS; his overall predicted score follows.

JLIB = —.82 + 2.46(3) + 1.81(4) + 1.59(5)


= 82 7,56% 724447.65
= 21.75
Another more conservative judge, whose scores are ABOR = 2, CAPP = 1;
and CENS = 3, would have this overall predicted score:
Additional Aspects of Correlation and Regression Analysis » 527

JUIB = 8222 2462) 4 1 Sicyt 10593)


Se Oo el 4
= 10.68
Remember that these are the scores predicted from the regression for-
mula. The actual overall scores assigned to the latter two judges have been
coded into the system. We expect our predictions to come close to the
actual scores, though they rarely will be exactly identical. R’ will tell us
approximately how close overall the predicted values come to the values
actually observed.
Let us return to the subject of the partial regression slopes in the
formula. Earlier we discussed the meaning of partial slopes in terms of a
visual example. We can now explain such slopes in terms of their impact
on prediction. We start with a two-variable example. Suppose we wanted
to predict JLIB from only one variable, ABOR. In that case, the regression
formula would be

JLIB = 10.63 + 2.39 ABOR


If ABOR = 1,
JLIB = 10.63 + 2.39(1)
= 10.63 + 2.39
= 13.02

If we increased ABOR by one unit, so that ABOR = 2, then

JLIB = 10.63 + 2.39(2)


= 10.63 + 4.78
a=) Eo?il

15.41— 13.02 =
Note that the difference between the two predictions is
2.39, the slope of ABOR. Thus, # tells us the amount of change in the depen-
independent
dent variable, JLIB, generated by a one-unit increase in the
ABOR's slope been negative, so
variable, ABOR. Note that had the sign of
ABOR would have
that JLIB = 10.63 — 2.39 ABOR, then a one-unit increase in
s would be
caused JLIB to decrease by 2.39 units because the variable
inversely related.
r independent
Suppose we wanted to predict from ABOR plus anothe
variable, CAPP. Our formula would be

VIB = 2.28 2.55 ABORS 2.48 CAPP


528 << STATISTICS FOR THE SOCIAL SCIENCES

The slope on ABOR, 2.55, is now a first-order partial regression slope, telling
as that if ABOR increased by one unit while CAPP remained unchanged, JLIB
would increase by 2.55 units. (Saying that CAPP remained unchanged is
another way of saying that we are “controlling for” that variable.) Also,
CAPP’s 6} of 2.48 tells us that if ABOR remained unchanged and CAPP
increased by one unit, JLIB would increase by 2.48 units.
Returning to our three-independent-variable formula,

JLIB = —.82 + 2.46 ABOR + 1.81 CAPP + 1.59 CENS

the slope of ABOR, 2.46, is a second-order partial regression slope that


shows the change in JLIB generated by a one-unit increase in ABOR when
both CAPP and CENS remain unchanged. The slope is a second-order par-
tial slope because we are controlling for two other variables.
Since the size of the slope measures the impact of the change on the
dependent variable, coming from a one-unit increase of the slope’s inde-
pendent variable, it would appear that the larger the slope, the more impor-
tant is that variable in explaining the dependent variable. Thus, ABOR is
more important than CAPP and CAPP is a bit more important than CENS.
Indeed, in this problem, that turns out to be the case. Generally, though, we
cannot reach this conclusion by examining the regression slopes because
the independent variables are usually measured along scales of differing
ranges and dispersions. Recall that in this problem, each independent vari-
able is measured on a scale ranging from 1 to 5. Suppose instead that ABOR
was scored from a low of 10 to a high of 50. Then its slope would have been
.246 instead of 2.46, and it would falsely appear to be less important than
CAPP or CENS. The easiest way to avoid this pitfall is to use a scale of iden-
tical range and standard deviation for each of the independent variables. If
this is not possible, we use the alternative approach presented next.

THE STANDARDIZED PARTIAL REGRESSION SLOPE


To compensate for differences in measurement, when we want to deter-
mine the relative importance of our independent variables, we calculate a
new measure known as a beta coefficient or a beta weight. Beta (f) is a
standardized partial regression slope. By standardized, we mean con-
verted into standard deviation units of the dependent variable, much in the
same vein as changing values of x into standard z scores. Beta tells us by how
many of its standard deviation units the dependent variable will change when
the variable associated with that beta weight increases by one of its standard
deviation units and the other independent variables remain unchanged.
Additional Aspects of Correlation and Regression Analysis » 529

Beta coefficient or beta weight A measure that is calculated when the relative
importance of the independent variables needs to be determined.

Standardized partial regression slope A slope expressed not in the original units
used but in standard deviation units of the dependent variable.

To get each beta, we multiply its b by its standard deviation and divide
by the standard deviation of the dependent variable.

SABOR
for ABOR, BABOR te bABOR x
STLIB
x

Ss
for CAPP, Bcapp — bcapp x oa
SJLIB

and so on.
In this case, for predicting JLIB,

Oran oh

Bree =o

CENS — .26

Thus, if ABOR increases by one of its standard deviation units and both
CAPP and CENS remain unchanged, JLIB will increase by .77 of one of its
standard deviation units. This is much larger a change than would result
from a standard deviation increase in either of the other two variables.
set.
As a further illustration of the value of betas, note the following data

x,= X= CS
74 6.7 ite
68 6.2 69

7.8 6.9 80
6.6 6.2 65
x, and .x,, the 6 for
When a regression is run on the above to predict x, from
x, from a scale run-
x, is 0.35, and the b forx, is 6.40. However, if we change
to 100 (75, 69, 81, etc.), the
ning from 0 to 1.00 to a scale running from 0
x, changes from 6.40 to 0.064.
b for x, is unchanged at 0.35, but the b for
, b, > b,. However, the beta
In the first example, b, < 0;, and in the second
530 << STATISTICS FOR THE SOCIAL SCIENCES

coefficients will be the same for both examples, 0.37 for x, and 0.66 for x,.
Regardless of the scale used for x,, it has about twice the impact on the
dependent variable as x,. Thus, while a comparison of the unstandardized
regression slopes varies with the scales of measurement of the independent
variables, this is not true with the standardized partial regression slopes. The
betas are unchanged.
Beta will always have the same sign as its b, so if b is negative, beta will
tell us by how many standard deviations the dependent variable will
decrease when that independent variable increases by a standard deviation.
Also, we normally never use the betas in a prediction equation. Accordingly,
we use the partial bs in the formula so that we are predicting the indepen-
dent variable value in its original scale. We would rarely want to make pre-
dictions in standard deviation units, which use of the betas would do. The
betas are used to compare the relative impact of each independent variable
on the dependent variable. The bs are used for prediction. Consequently,
each coefficient has its own role to play.

USING A REGRESSION PRINTOUT

In Table 14.2, you will see a sample multiple regression printout using
SAS’s regression procedure. (Setup instructions for the SAS procedures
will be discussed shortly.) In this case, all five of the original indepen-
dent variables in our problem were retained. The printout begins with an
ANOVA for the overall statistical significance of the regression equation.
The F of 72.606 has a probability of .0001—very significant. Below the
ANOVA is additional information. Note in the center that an R-SQUARE
value of 0.8561 is given. Thus, 85.61% of the variation in JLIB can be
accounted for by this regression with five independent variables. Below
the R-SQUARE, note the ADJ R-SQ, which stands for adjusted R-Square.
The adjustment does for R what 7 — 1 in the denominator does for the
standard deviation. Just as with S, an R from sample data tends to overstate
the R in the population. With the adjustment, we may estimate that in the
population as a whole (all similar judges), 84.43% of the variation in JLIB
can be accounted for by this model.
LBBLLLLLLLLLLAMLLAMRAPLPLPA LLL LLL LAELIA LAL AED ALN ELLE COLE CN TTT een etsmeNte

Adjusted R-Square An estimate of R-Square in the population from which the


sample was drawn.

In the next lower part of the printout, you will find the information
needed to construct the regression formula. (Why doesn’t this procedure
just generate the formula?) The first column on the left is entitled VARIABLE.
Additional Aspects of Correlation and Regression Analysis jp 531

Table 14.2 The SAS Regression Procedure

SAS
DEP VARIABLE: JLIB ANALYSIS OF VARIANCE

SUM OF MEAN
SOURCE DF SQUARES SQUARE F VALUE PROB >F
MODEL 5 3010.44928 602.08986 72.606 0.0001
ERROR 61 505.84923 8.29261031
C TOTAL 66 3516.29851
ROOT MSE 2.879689 R-SQUARE 0.8561
DEP MEAN 19.41791 ADJ R-SQ 0.8443
CV 14.83007

PARAMETER ESTIMATES
PARAMETER STANDARD T FOR Ho;
VARIABLE DF ESTIMATE ERROR — PARAMETER=0 PROB > |T|
INTERCEP 1 ~0.96098181 1.37008962 ~0.701 0.4857
CENS 1 1.40981955 0.34490030 4.088 0.0001
CAPP 1 1.45120569 0.32446883 4.473 0.0001
CONS 1 0.97003442 0.41702892 2.326 0.0234
RDIE 1 ~0.02914917 0.01595688 ~1,827 0.0726
ABOR 1 2.40255343 0.15804215 15.202 0.0001

STANDARDIZED
VARIABLE DF ESTIMATE
INTERCEP 1 0
CENS 1 0.23138398
CAPP 1 0.27146741
CONS 1 0.13229489
RDIE 1 —0.09153555
ABOR 1 0.75883123

Below the title is INTERCEP meaning, of course, the intercept @) for the vari-
able JLIB. Below INTERCEP are the names of the variables in the order in which
the program was requested to enter them.
DF,
To the right of the VARIABLE column is a degrees-of-freedom column,
ion
and to its right, PARAMETER ESTIMATE. In this latter column is the informat
intercept , —0.9609 8181, and
we need for our equation. The first number is the
variable entered,
the number under it is the partial regression slope for the first
532 <4 STATISTICS FOR THE SOCIAL SCIENCES

CENS, which is 1.40981955. (From now on, for simplicity’s sake, we'll use
only the first two decimal places, and to ease your reading of the printout,
we will not round the entries.) Going down this column, we can build our
regression formula:

JLIB = —0.96 + 1.40 CENS + 1.45 CAPP + 0.97 CONS


— 0.02 RDIE + 2.40 ABOR

After the intercept of —0.96, we could have entered the variables in any
order that we preferred: JLIB = —0.96 + 2.40 ABOR + 1.45 CAPP ... and so
on. Note that below these numbers on the printout is a column titled
STANDARDIZED ESTIMATE. These are the beta coefficients. (The inter-
cept, of course, has no beta.) Note that both the 4s and betas differ from
those used earlier in this chapter because now two new variables have
been added.
Moving back up to the PARAMETER ESTIMATE column and moving two
columns to the right, you’ll see T FOR HO; PARAMETER = 0 and to its right
a PROB > |T| column. In addition to the overall F already done, SAS tests
each component of the equation for significance. The first ¢ value, —0.701,
tests the null hypothesis @,,.utation = 0. Since p= 0.4857, we cannot reject A).
This is OK because the printed intercept is small to begin with, —0.96. For
the population’s regression formula, we could substitute 0 for —0.96. (Make
no substitution, of course, ifp < .05.) The rest of the ¢ tests are for the null
hypothesis: 8... uiation = 9. With the exception of RDIE, all ¢s are significant and
their respective variables retained in the equation.
Since the ¢ for RDIE is not significant, we must assume D..nulation = 9: If
that is the case, remember that 7,,,,,.iation = 0- It would be preferable to rerun
this procedure without using RDIE at all, as shown in Table 14.3. In addition,
the independent variables were entered in decreasing order of the size of
their slopes. Note that just as when we add a variable the coefficients
change, the same thing happens when we delete one. Our new equation
becomes

JLIB = -1.57 + 2.44 ABOR + 1.55 CAPP + 1.44 CENS


+ 0.91 CONS

Since the intercept is not significantly different from zero, we could delete it
as well. (However, many programs do not test the intercept for significance,
so it is rarely deleted in practice.)

JLIB = 2.44 ABOR + 1.55 CAPP + 1.44 CENS + 0.91 CONS


Additional Aspects of Correlation and Regression Analysis p 533

Table 14.3 Regression Without RDIE and Independent Variables Reordered

SAS
DEP VARIABLE: JLIB ANALYSIS OF VARIANCE

SUM OF MEAN
SOURCE DF SQUARES SQUARE F VALUE PROB > F
MODEL 4 2982.77683 745.69421 86.656 0.0001
ERROR 62 533.52167 8.60518829
C TOTAL 66 3516.29851
ROOT MSE 2.93346 R-SQUARE 0.8483
DEP MEAN 19.41791 ADJ R-SQ 0.8385
CV, 15.10698
PARAMETER ESTIMATES
PARAMETER STANDARD T FOR HO:
VARIABLE DF ESTIMATE ERROR — PARAMETER=0 PROB >|T |
INTERCEP 1 —1.57963264 1.35236320 -1.168 0.2473
ABOR 1 2.44618610 0.15914391 i571 0.0001
CAPP 1 1.55217744 0.32569620 4.766 0.0001
CENS 1 1.44292853 0.35085499 4.113 0.0001
CONS 1 0.91434287 0.42367918 2.158 0.0348
STANDARDIZED
VARIABLE DF ESTIMATE
INTERCEP 1 0
ABOR 1 0.77261233
CAPP 1 0.29035552
CENS 1 0.23681793
CONS 1 0.12469958

STEPWISE MULTIPLE REGRESSION

In the above applications, it was up to the researcher to specify the


independent variables for the regression model. A number of techniques
subsumed under the title of stepwise multiple regression allow the
computer to determine which independent variables to enter and when to
enter them. In theory, we are telling the computer to first find the “best”
independent variable for predicting the dependent variable. The computer,
in Step 1 below, finds that “best” variable among a menu of variables and
gives us a simple regression for predicting the dependent variable from the
selected independent variable.
534 STATISTICS FOR THE SOCIAL SCIENCES

Stepwise multiple regression A procedure in which each independent variable is


added to the regression in a separate step. The order of entry is based on criteria
selected by the researcher.

Then the computer finds the “second best” independent variable and,
in Step 2, adds that variable to the model. Now we have two independent
variables in a regression. Then the computer finds the “third best” indepen-
dent variable and adds it to the equation. It continues doing this until it runs
out of independent variables, or out of time, or other preselected criteria
are met for shutting down the process.
There are a number of techniques for finding the “best,” “second best,”
and so on, independent variables. One common way is as follows:

Step 1: Do a matrix of simple correlations and enter the variable most


highly correlated with the dependent variable.

Step 2: Do a matrix of first-order partial correlations (controlling for the


variable selected in Step 1). Find in this matrix the variable most highly
correlated with the dependent variable and use this new variable along
with the one selected in Step 1 to create an equation for predicting the
dependent variable.

Step 3: Find the variable whose second-order partial correlation (con-


trolling for the two previously selected variables) correlates the highest
and enter it in Step 3 in a new regression equation.

At each step, the program generates an R or an R’. The first one (hope-
fully) will be large. The one in the second step will be even larger than the
first since two variables should account for more variation than just one.
However, with each subsequent step, the increment in the size of the coef-
ficient will get smaller, as each latter variable tends to bring about only a
small increase in explained variation. Watching to see where R° “tops off,”
together with an analysis of the significance of the regression slopes, enables
the researcher to decide which model to select.
In Table 14.4, we see the results of the SAS PROC STEPWISE procedure
as applied to our judicial scaling problem. Each step resembles, in general,
the earlier format in PROC REG.
At the beginning of each step, the printout shows the step number, the
independent variable entered, and R-SQUARE. (STEPWISE does not find the
adjusted R* and does not calculate beta coefficients. However, once you
have selected the variables for your final model, you could also run PROC
REG, as we did previously, to get this information.)
Additional Aspects of Correlation and Regression Analysis > 535

Table 14.4 SAS Stepwise Regression

SAS
STEPWISE REGRESSION PROCEDURE FOR DEPENDENT VARIABLE JLIB
NOTE: SLENTRY AND SLSTAY HAVE BEEN SET TO
.-15 FOR THE STEPWISE TECHNIQUE.

R SQUARE = 0.57096438
STEP 1 VARIABLE ABOR ENTERED C(P) = 118.92309202
DF SUM OF SQUARES MEAN SQUARE FF PROB > F
REGRESSION if 2007.68119830 2007.6811983 86.50 0.0001
ERROR 65 1508.61730917 23.2094971
TOTAL 66 3516.29850746

B VALUE STD ERROR TXRERTICSS: F PROB > F


INTERCEPT = 10.63390350
ABOR 2.39239214 Oy 2 2007.6811983 86.50 0.0001
BOUNDS ON CONDITION NUMBER: tl 1

R SQUARE = 0.78386266
STEP 2 VARIABLE CAPP ENTERED C(P) = 30.64827058

DF SUM OF SQUARES MEAN SQUARE PF PROB > F

REGRESSION 2, 2756.29511365 1378.1475568 116.05 0.0001


ERROR 64 760.00339381 11.8750530

TOTAL 66 3516.29850746

B VALUE STD ERROR TYPE II SS F PROB >F

INTERCEPT 2.283 13643


CAPP 2.48235217 0.31264553 748.6139154 63.04 0.0001
ABOR De SSIS Tes 0.18516922 2265.8458299 190.81 0.0001

BOUNDS ON CONDITION NUMBER: 1.01282,


4.051278

R SQUARE = 0.83687406
STEP 3 VARIABLE CENS ENTERED C(P) = 10.16995879

DF SUM OF SQUARES MEAN SQUARE F PROB > F

REGRESSION 3 294269899378 980.89966459 107.73 0.0001


ERROR 63 573.59951368 9.10475419

TOTAL 66 3516.29850746 z

(Continued)
536 << STATISTICS FOR THE SOCIAL SCIENCES

Table 14.4 (Continued)

B VALUE STD ERROR TYPE II SS F PROB > F


INTERCEPT —0.82773845
CENS 1.59823760 0.35322219 186.4038801 20.47 0.0001
CAPP 1.81212925 0.31126325 308.5964051 33.89 0.0001
ABOR 2.46070774 0.16355183 2060.9976536 226.36 0.0001
BOUNDS ON CONDITION NUMBER: 1.309335,
10.91348

R SQUARE = 0.84827179
STEP 4 VARIABLE CONS ENTERED C(P) = 7.33700054
DF SUM OF SQUARES — MEAN SQUARE F PROB > F
REGRESSION 4 2982.77683321 745.69420830 86.66 0.0001
ERROR 62 533.52167426 8.60518829
TOTAL 66 3516.29850746
B VALUE STD ERROR TYPE II SS F PROB >F
INTERCEPT -1.57963264
CENS 1.44292853 0.35085499 145.5441367 16.91 0.0001
CAPP 1.55217744 0.32569620 195.4419285 22.71 0.0001
CONS 0.91434287 0.42367918 40.0778394 4.66 0.0348
ABOR 244618610 0.15914391 2033.1026734 226.36 0.0001
BOUNDS ON CONDITION NUMBER: 1.516799,
21.0738

R SQUARE = 0.85614156
STEP 5 VARIABLE RDIE ENTERED C(P) = 6.00000000
DF SUM OF SQUARES — MEAN SQUARE F PROB > F
REGRESSION 5 3010.44927833 602.08985567 72.61 0.0001
ERROR 61 505.84922913 8.29261031
TOTAL 66 3516.29850746
B VALUE STD ERROR TYPE II SS F PROB >F
INTERCEPT -0,96098181
CENS 1.40981955 0.34490030 138.5578569 16.71 0.0001
CAPP 1.45120569 032446883 165.8835 156 20.00 0.0001
CONS 0.97003442 0.41702892 44.8676379 5.41 0.0234
RDIE -0.02914917 0.01595688 27.6724451 3.34 0.0726
ABOR 2.40255343 0.15804215 1916.4236580 231.10 0.0001
BOUNDS ON CONDITION NUMBER: 1.562133,
32.06837
NO OTHER VARIABLES MET THE 0.1500 SIGNIFICANCE LEVEL FOR ENTRY
Additional Aspects of Correlation and Regression Analysis j» 537

Below this information, an F is generated for that model and its


probability given. Below the analysis of variance in the column titled B
VALUE, you will find the same information as was in the column PARAMETER
ESTIMATES in the REG procedure done earlier.
In Step 1, therefore, we see that RSQUARE = 0.57 (we again report only
to two decimal places). This 0.57 in the first step is the simple 7~ between
JLB and ABOR. The intercept is 10.63, and the slope of ABOR is 2.39. Also
note that in each step, each slope (but not the intercept) is tested for
significance. STEPWISE does this with F rather than ¢, but the probabilities
are identical to those in REG. We construct from the data the first equation.

JLIB = 10.63 + 2.39 ABOR

For Step 2, CAPP is entered, and R’ increases from .57 in Step 1 to .78.
The new intercept is 2.28, the slope for CAPP is 2.48, and the new slope
for ABOR is 2.55. (In later steps, you will see that in the list of indepen-
dent variables, the variables are not listed in the same order as they were
brought into the model; rather, they are listed in the order in which they
were specified when the run was set up.)
Following is a summary of the regression equations for each step and
with the independent variables listed in the order in which they were
brought in.

Step Re Equation
1 57 JLIB = 10.63 + 2.39 ABOR
2 Ake JLIB = 2.28 + 2.55 ABOR + 2.48 CAPP
3 heels, JLIB = -0.82 + 2.46 ABOR + 1.81 CAPP + 1.59 CENS
4 84 JLIB = -1.57 + 2.44 ABOR + 1.55 CAPP + 1.44 CENS
+ 0.91 CONS
B) me) JLIB = -0.96 + 2.40 ABOR + 1.45 CAPP + 1.40 CENS
+ 0.91 CONS — 0.02 RDIE

As before, RDIE in Step 5 is not statistically significant, p = .0726. In addi-


tion, CONS, while significant in both Step 4 (p = .0348) and Step 5 (p= .0234),
has a much higher probability of error than the other three independent vari-
ables. Also, R? reaches .83 in Step 3 and grows very little in the subsequent
steps. For these reasons, we select the model in Step 3 as our “working
model” for further analysis. This is why this model was given to you earlier in
the chapter as the “best” model for our example.
A word of caution here: Because the stepwise procedure is based on
the
statistical considerations alone and requires no theoretical justification for
538 @ STATISTICS FOR THE SOCIAL SCIENCES

independent variables it includes or excludes from entry, we know what was


and was not included but not why. Perhaps consumer protection and right-
to-die issues have not been under judicial consideration long enough for legal
scholars to see them ina liberal/conservative continuum. Perhaps these issues
are judged on very situation-specific considerations. Whatever the reasons,
this procedure may raise more questions than it answers, a weakness of many
techniques where many variables are manipulated simultaneously.
At the same time, the questions raised may in turn give rise to theory.
How many cases in a given issue area must be considered before the area
itself becomes subject to judicial ideology? What other crosscutting values
may be present in the judges’ minds? The stepwise procedure is like a
fishing expedition. We cast our nets into waters that are murky because we
have little theory to guide us. We collect whatever our nets entangle, not
knowing until later if the fish we caught are even edible. Still, it is fun to
go fishing, and sometimes we make a good catch. Many scholars, however,
are mistrustful of “fishing expeditions.” They would rather the researcher
enter in the variables, basing the selection of the variables to be entered on
a theoretical foundation rather than the “whims” of the computer.

COMPUTER APPLICATIONS

Partial Correlations—SPSS

Imagine a study of 25 adult repeat offenders imprisoned for homicide.


You wish to predict the number of years each offender will serve, SEN-
TENCE, based on that person’s number of arrests for various crimes prior to
the homicide. The crimes are the following: ARSON, ASSAULT, BURGLARY,
RAPE, ROBBERY. The data are entered in Table 14.5. The first person entered
has been arrested twice for arson, four times for assault, and three times
each for burglary, rape, and robbery; that prisoner is now serving a 10-year
sentence on the homicide conviction, (Perhaps he is the judge’s cousin.)
We wish to examine for all our prisoners the correlation between ARSON
and SENTENCE first as a zero-order correlation and then as first-, second-,
and third-order partial correlations, controlling first for RAPE, then RAPE and
ROBBERY, and finally controlling for RAPE, ROBBERY, and ASSAULT.
Since we now have six variables to keep track of, we have replaced
VAROO001, VAROOO002, and so on with variable names. Remember that to do
this, we click on variable view at the bottom of the screen, left click in each
name box, and type in the actual name. Click back to data view, and the new
names appear as in Table 14.5. (Please note that the reproductions of the
printouts in this chapter may differ slightly from the output you would get
Additional Aspects of Correlation and Regression Analysis j» 539

Table 14.5

ARSON ASSAULT BURGLARY RAPE ROBBERY SENTENCE

il 2.00 4.00 3.00 3.00 3.00 10.00


2 OO 2.00 2.00 2.00 2.00 12.00
2 1.00 2.00 1.00 OO 4.00 12.00
4 2.00 2.00 ‘ 4.00 00 1.00 13.00
>) 1.00 4.00 3.00 1.00 1.00 22.00
6 2.00 1.00 2.00 2.00 5.00 16.00
vi OO 4.00 3.00 OO 3.00 20.00
8 OO 2.00 1.00 3.00 4.00 14.00
9 00 2.00 3.00 3.00 3.00 16.00
10 3.00 1.00 2.00 1.00 2.00 16.00
11 1.00 4.00 3.00 1.00 4.00 20.00
ie, 4.00 2.00 1.00 2.00 3.00 17.00
16} OO 5.00 3.00 1.00 1.00 16.00
14 1.00 5100 3.00 2.00 4.00 19.00
Is OO 5.00 2.00 1.00 5.00 26.00
16 1.00 5.00 2.00 1.00 4.00 23.00
Wy 1.00 1.00 1.00 2.00 4.00 12.00
18 3.00 2.00 2.00 2.00 3.00 17.00
iS) OO 5.00 4.00 OO 5.00 36.00
20 OO 2.00 1.00 2.00 3.00 13.00
Zit .OO 2.00 2.00 .OO 5.00 17.00
Daje 2.00 4.00 4.00 OO 5.00 26.00
2B 4.00 4.00 3.00 2.00 2.00 19.00
24 OO 4.00 3.00 1.00 5.00 25.00
7S OO 1.00 1.00 2.00 5.00 14.00

replicating the run. This is for editorial or stylistic reasons. For instance, the
y in burglary and the last e in sentence actually appear in the line below
where they appear in Table 14.5. SPSS only allocates seven spaces per line
for the title. Here, we have opted to keep our words whole.)
As in Chapter 13, we get the zero-order coefficient by clicking

Analyze
Correlate

Bivariate

and moving ARSON and SENTENCE into the Variables box. Then we click ok.
540 << STATISTICS FOR THE SOCIAL SCIENCES

To get our partials, each time we will click

Analyze
Correlate

Partial

and move ARSON and SENTENCE into the Variables box. Then we move
RAPE into the box indicating the control variable and click ok. We repeat
this procedure again but bring ROBBERY in under RAPE in our control vari-
able list and click ok. We repeat the same procedure once again, adding
ASSAULT to ROBBERY and RAPE in the control variable list before clicking
ok. Ali results are summarized in Table 14.6. The full matrix of zero-order
correlations run on SPSS is presented in Table 14.7.

Partial Correlations—Other Programs

The SAS ANALYST procedure does not yet have a subroutine for partial
correlation. It is possible to use the regular SAS programming language to
get partials, but since the latter has not been covered in this edition of the
book, we will not discuss it here. Likewise, Excel’s statistics add-on does nei-
ther partial correlations nor stepwise multiple regression. The procedure
for finding zero-order correlations in both SAS and Excel was presented in
the previous chapter.

Multiple Regression—SPSS

Starting with the data list in Table 14.5, click the following:

Analyze
Regression

Linear

Put SENTENCE in the dependent variables box and the other five variables
in the independent variables box. Leave the methods button set on Enter:
The independent variables will be listed on the output screen in the same
order that you entered them into the independent variables box. Click ok.
The output is found in Table 14.8.
Additional Aspects of Correlation and Regression Analysis j» 541

Table 14.6 ~Correlation—SPSS


Correlations

Correlations

ARSON SENTENCE

ARSON Pearson Correlation 1 —.132


Sig. (2-tailed) 530
N 25 25
SENTENCE Pearson Correlation —.132 1
Sig. (2-tailed) 530
N 25 25

Partial Corr

Correlations

Control Variables ARSON SENTENCE

RAPE ARSON Correlation 1.000 —.082


Significance (2-tailed) 704
df 0 2p.

SENTENCE Correlation —.082 1.000


Significance (2-tailed) ./04
df 22 O

Correlations

Control Variables ARSON SENTENCE

RAPE & ARSON Correlation 1.000 031


ROBBERY Significance (2-tailed) 889
df 0 21

SENTENCE Correlation 031 1.000


Significance (2-tailed) 889
df PAL 0

Partial Corr
Correlations

Control Variables ARSON SENTENCE

ARSON Correlation 1.000 194


RAPE &
Significance (2-tailed) 388
ROBBERY &
df 0 20
ASSAULT
SENTENCE Correlation 194 1.000
Significance (2-tailed) 388
df 20 0
542 € STATISTICS FOR THE SOCIAL SCIENCES

Table 14.7. Zero-Order Correlations—SPSS

Correlations

ARSON ASSAULT BURGLARY RAPE


a a a 8... ——————E—

ARSON Pearson 1 —176 .030 126


Correlation : 400 888 548
Sig. (2-tailed) 25 25 Wp, 25
N

ASSAULT Pearson —.176 1 Be fo ole —335


Correlation 400 : 002 LOZ
Sig. (2-tailed) 25 25 25 25
N

BURGLARY Pearson .030 Door A 1 —,389


Correlation 888 002 .055
Sig. (2-tailed) 25 25 25 aS
N

RAPE Pearson 126 = 3)05) —,389 1


Correlation 548 .102 .055
Sig. (2-tailed) 25 25 25 aS
N

ROBBERY Pearson —.291 —.003 —.184 —.091


Correlation lls! 990 379 664
Sig. (2-tailed) 25 25 25 25
N

SENTENCE Pearson —,132 .670** —.519** -


Correlation 50) 000 008 .481*
Sig. (2-tailed) 25 25 25 015
N 25

Correlations

ROBBERY SENTENCE
ARSON Pearson —.291 —.132
Correlation sls 530
Sig. (2-tailed) 25 25
N

ASSAULT Pearson —.003 .670**


Correlation 990 .000
Sig. (2-tailed) 25 25
N
Additional Aspects of Correlation and Regression Analysis je 543

BURGLARY Pearson —.184 JSulligee


Correlation 379 008
Sig. (2-tailed) Pass 25
N

RAPE Pearson = O91) —.481


Correlation 664 O15
Sig. (2-tailed) 25 JAS,
N

ROBBERY Pearson {| 380


Correlation 061
Sig. (2-tailed) aS oS)
N

SENTENCE Pearson 380 1


Correlation 061 f
Sig. (2-tailed) 25, ip
N

*Correlation is significant at the .05 level (2-tailed). **Correlation is significant at the .01 level
(2-tailed).

If you compare this table to the earlier SAS format in Table 14.2
(a different data set), SPSS lists the intercept [Constant] first and then
the variables with their slopes in the B column. SAS also presents the
intercepts first, to the left of INTERCEP in the PARAMETER ESTIMATE
column. The other entries in that column are the slopes. The SAS regres-
sion results equivalent to the SPSS run shown in Table 14.8 appear in
Table 14.10.

Multiple Regression—SAS

To get the output presented in Tables 14.9 and 14.10, do the following:

Solutions
Analysis

Analyst
STATISTICS FOR THE SOCIAL SCIENCES

Table 14.8 Regression—SPSS

Variables Entered/Removed?

Variables Variables
Model Entered Removed Method

1 ROBBERY,
ASSUALT,
ARSON, Enter
RAPE,
BURGLARY*

a. All requested variables entered.


b. Dependent variable: SENTENCE.

Model Summary

Adjusted R Std. Error of


Model R R Square Square the Estimate

1 834" 695 615 3.6398

a. Predictors: (Constant), ROBBERY, ASSAULT, ARSON, RAPE, BURGLARY.

ANOVA?

Sum of Mean
Model Squares af Square F Sig.

1 Regression 5/3242 5 114.048 8.054 0008


Residual 251.018 iB 13.248
Total 824.960 24

a. Predictors: (Constant), ROBBERY, ASSAULT, ARSON, RAPE, BURGLARY.


b. Dependent variable: SENTENCE.

Coefficients*

Unstandardized Standardized
Coefficients Coefficients

Model B Std. Error Beta t Sig.

il (Constant) Sito 3.971 894 382


ARSON 443 .613 .098 ae 479
ASSAULT 2.009 .673 484 2.984 008
BURGLARY 1.385 pele) 299 L395 yO
RAPE —1,182 833 ik —-1.419 BE
ROBBERY 1.880 589 435 eae. I 005
a. Dependent variable: SENTENCE.
Additional Aspects of Correlation and Regression Analysis j» 545

Table 14.9 Correlations—SAS

The CORR Procedure

6 Variables: A B G D E F

Simple Statistics

Variable N Mean Std Dev Sum Minimum Maximum

A DS 1.12000 1.30128 28.00000 0 4.00000


B 25 2.92000 1.41185 73.00000 1.00000 5.00000
Cc 25 2.36000 0.99499 59.00000 1.00000 4.00000
D 25 1.36000 0.99499 34.00000 0 3.00000
E 25 3.44000 1.35647 86.00000 1.00000 5.00000
F De) 18.04000 5.86288 451.00000 10.00000 36.00000

Pearson Correlation Coefficients, N = 25


Prob > |r| under HO: Rho = 0

A B G D E F

A 1.00000 =O, 0.02961 0.12615 —0.29082 SUS 175


0.4001 0.8883 0.5479 0.1584 0.5302
B 0.17599 1.00000 0.58491 (sien 5y —0.00261 0.66989
0.4001 0.0021 0.1021 0.9901 0.0002
€ 0.02961 0.58491 1.00000 —0.38889 —0.18400 0.51884
0.8883 0.0021 0.0547 0.3786 0.0079
D 0.12615 =O 5407 —0.38889 1.00000 =0,09136 —0.48113
0.5479 0.1021 0.0547 0.6640 0.0149
E —0.29082 —0.00261 0.18400 —0,09136 1.00000 0.38016
0.1584 0.9901 0.3786 0.6640 0.0609
F moO Rlol Wee, 0.66989 0.51884 —0.48113 0.38016 1.00000
0.5302 0.0002 0.0079 0.0149 0.0609

Enter the data just as in Table 14.5. Then,

Statistics
Descriptive

Correlations

The correlation matrix appears in Table 14.9. Now, for the regression, click
back to the data entry page, and click
546 STATISTICS FOR THE SOCIAL SCIENCES

Statistics
Regression

Linear

Highlight A, B, C, and D and move them into the Explanatory box. Highlight
F and move it into the Independent box. Click ok. The full model regression
appears in Table 14.10.

Table 14.10 Multiple Regression—SAS

The REG Procedure


Model: MODEL 1
Dependent Variable: F

Number of Observations Read Ss.


Number of Observations Used 5)

Analysis of Variance

Source DF Sum of Squares Mean Square F Value [eres Ii

Model 5 573.24243 114.64849 8.65 0.0002


Error 19 25h. 74757 13.24829
Corrected Total 24 824.96000
Root MSE 3.63982 R- 0.6949
Dependent 18.04000 Square 0.6146
Mean 20.17639 Adj R-Sq
Coeff Var

Parameter Estimates

Squared Squared
Parameter Standard Partial Corr Type I Partial
Variable DF Estimate Error tValue Pr>|t| Corr Type 1 Corr Type Il

Intercept 1 a) sy 3.97121 0.89 0.3824


A | 0.44254 0.61265 0.72 0.4789 0.01735 0.02673
B | 2.00881 0.67322 2.98 0.0076 0.43921 0.31909
3 1 1.38485 0.99303 1,39 0.1792 0.04714 0.09286
D ] —1.18162 0.83289 —1.42 Omy22 0.10770 0.09578
E 1 1.87974 0.58930 a.19 0.0048 0.34875 0.34875
Additional Aspects of Correlation and Regression Analysis j 547

Multiple Regression—Excel

Key in the crime data, as before. To get the correlation matrix, click

Tools

Data Analysis

Correlation

Click ok and highlight the input range, as before. It should read $A$1:$F $25.
Click ok.
For the regression run, click

Tools
Data Analysis
Regression

Click ok and highlight the input x range (column F). In that box, it should
read $F$1:$F$25. Click to the imput x range box. Then highlight the rest of
the data and click it into the box. It should read $A$1:$E$25. Click ok. Both
the correlation and regression output are displayed in Table 14.11.

Stepwise Multiple Regression—SPSS

For stepwise multiple regression, call the program as before:

Analyze
Regression
Linear

As before, put SENTENCE in the dependent variable box and the other vari-
ables in the independent variables box. Now go to the method button where
it says Enter; click and replace that word with Stepwise.
One other step is needed. All stepwise procedures have criteria for
accepting or deleting independent variables from the model. The criteria
may be based on the level of significance of F in the analysis of variance
that accompanies each variable entry or in the amount a new variable
would increase R*. These limits help keep the regression from growing
large with variables that do little to enhance the predictability of the
dependent variable. As its default option, SPSS requires the probability of
548 <4 STATISTICS FOR THE SOCIAL SCIENCES

Table 14.11 Correlation and Regression—Excel

Column 1 Column 2 Column 3 Column 4 Column5 Column 6

Column 1 1
Column 2 175079) 1
Column 3 0.029607 0.584909 1
Column 4 0.12615 —.334573 —.388889 1
Column 5 —.290817 —.002611 —.183996 —.091381 1
Column6 —-0.13173 0.669886 0.518843 —0.48113 0.38016 1

SUMMARY OUTPUT

Regression Statistics

Multiple R Wels eisey


R Square 0.694873
Adjusted R 0.614576
Standard E 3.63982
Observation 25

ANOVA

df SS MS Significance F

Regression 5 573.2424 114.6485 8.653831 0.000207


Residual tS) 251.7176 13.24829
Total 24 824.96

Standard Lower Upper Lower Upper


Coefficient —Error t Stat P-value 95% 95% 95.0% 95.0%

Intercept 3.551066 3.971212 0.894202 0.382401 -.760776 11.86291 -4.760776 11.86291


X Variable 0.442536 0.612648 0.722332 0.478887 -.839752 1.724824 -0.839752 1.724824
X Variable 2.008814 0.673216 2.983908 0.007629 0.599757 3.417871 0.599757 3.417871
X Variable 1.384855 0.993025 1.394581 0.179235 -.693571 3.463281 -—0.693571 3.463281
X Variable —1.181618 0.832894 -—.418689 0.172185 -.924884 0.561649 -—2.924884 0.561649
X Variable 1.879738 0.5893 3.189783 0.004823 0.64632 3.113157 0.64632 Sno57/
Additional Aspects of Correlation and Regression Analysis j» 549

F to be .05 or less to enter a variable and .10 or more to remove a variable.


If these values are retained in our sample problem, only ASSAULT and
ROBBERY would be brought into the model. To get all the independent
variables in, click the options button and change the entry and exit values
to .50 and .90, respectively. (Other values may work as well.) Press con-
linue to get to the main Dialog box and then click ok. The results appear
in Table 14.12. ‘

Table 14.12 Stepwise Regression—SPSS

Variables Entered/Removed*

Variables Variables
Model Entered Removed — Method

1 ASSAULT Stepwise (Criteria: Probability-of-F-to-enter


< .500, Probability-of-F-to-remove > .900).

2 ROBBERY Stepwise (Criteria: Probability-of-F-to-enter


< 500, Probability-ofF-to-remove = .900).

5) BURGLARY Stepwise (Criteria: Probability-of-F-to-enter


< .500, Probability-of-F-to-remove = .900).

4 RAPE Stepwise (Criteria: Probability-ofF-to-enter


< .500, Probability-of-F-to-remove = .900).

5 ARSON Stepwise (Criteria: Probability-of-F-to-enter


< .500, Probability-of-F-to-remove = .900).

a. Dependent variable: SENTENCE.

Model Summary

Std. Error of the


Model R R Square Adjusted R Square Estimate

1 .670* 449 425 4.4466


2 Ta 99> 558 3.8989
3 810° 657 .607 3.6751
+ B29" .686 624 3.5960
5 834° 695 615 3.0398

a. Predictors: (Constant), ASSAULT.


b. Predictors: (Constant): ASSAULT, ROBBERY.
c. Predictors: (Constant): ASSAULT, ROBBERY, BURGLARY.
d. Predictors: (Constant): ASSAULT, ROBBERY, BURGLARY, RAPE,
e. Predictors: (Constant): ASSAULT, ROBBERY, BURGLARY, RAPE, ARSON.

(Continued)
550 « STATISTICS FOR THE SOCIAL SCIENCES

Table 14.12 (Continued)

ANOVA‘

Sum of Mean
Model Squares df Squares F Sig.

1 Regression 370.198 370.198 18.723 0008


Residual 454.762 23 LZ,
Total 824.960 24

2 Regression 490.523 2 245.262 16.134 .000"


Residual 334.437 22 15.202
Total 824.960 24

5 Regression 541.631 3 180.544 13.382 .OO0S


Residual DES O29) 21 13.492
Total 824.960 24

4 Regression 566.330 4 141.582 10.949 .000"


Residual 258.630 20 W932
Total 824.960 24

5 Regression 573.242 5 114.648 8.654 OOO


Residual pli 19 13.248
Total 824.960 24

a. Predictors: (Constant), ASSAULT


b, Predictors: (Constant), ASSAULT, ROBBERY
c. Predictors: (Constant), ASSAULT, ROBBERY, BURGLARY
d. Predictors: (Constant), ASSAULT, ROBBERY, BURGLARY, RAPE
e. Predictors: (Constant), ASSAULT, ROBBERY, BURGLARY, RAPE, ARSON
f. Dependent variable: SENTENCE.

Coefficients*

Unstandardized Standardized
Coefficients Coefficients

Model B Std. Error Beta l Sig.


il (Constant) 9.917 2.007 4,774 OOO
ASSAULT Doz 043 .670 4.327 OOO

2 (Constant) 4,227 Dial Dee IRS Do slays:


ASSAULT 2.786 564 .671 4,942 OOO
ROBBERY 1.661 587 oy” Been) O10
Additional Aspects of Correlation and Regression Analysis > 551

3 (Constant) I 2S L992 409 .687


ASSAULT 2.022 .660 487 3.061 .006
ROBBERY 1.899 567 439 3.348 .003
BURGLARY 1.856 by58) Palsy 1.946 005
4 (Constant) 4.428 3.736 1.185 .250
ASSAULT LOLI 652 460 PAX: 008
ROBBERY eee7a 30 .410 3.146 005
BURGLARY 1.475 78 .250 ile lis 145
RAPE =1134 820 lle = {Lote 182
5 (Constant) 3051 RO 894, ew
ASSAULT 2.009 673 484 2.984 008
ROBBERY 1.880 589 435 3.190 005
BURGLARY 1.385 993 20> 1305) 179
RAPE —1.182 .833 = 2 Oil —1.419 ali72,
ARSON 443 G14 098 Py! A79

a. Dependent variable: SENTENCE.

Excluded Variables‘

Collinearity
Partial Statistics
Model Beta In t Sig. Correlation Tolerance

1 ARSON —.014° —.089 0 —.019 969


BURGLARY Uivee 1.012 p22 “PAiil .658
RAPE —.289" —1.852 ROW, ~.367 888
ROBBERY ghee 2.813 O10 514 1.000

Z ARSON Lage .756 458 163 884


BURGLARY oo. 1.946 .065 391 625
RAPE 2257" —1.833 081 =a) | .880

8) ARSON 083° SION aw85) ey 874


RAPE —.192° -1.382 .182 —.295 .809

4 ARSON 098" ee 479 .163 869

a. Predictors in the model: (Constant), ASSAULT.


b. Predictors in the model: (Constant), ASSAULT, ROBBERY.
c. Predictors in the model: (Constant), ASSAULT, ROBBERY, BURGLARY.
d. Predictors in the model: (Constant), ASSAULT, ROBBERY, BURGLARY, RAPE.
e. Dependent variable: SENTENCE.
552 << STATISTICS FOR THE SOCIAL SCIENCES

The independent variables are brought into the equation starting with
the one most highly correlated with the dependent variable, SENTENCE:
ASSAULT, ROBBERY, BURGLARY, RAPE, and ARSON. Note that under coeffi-
cients, in the last step (Step 5), the data are identical to the regression we
originally did (Table 14.8).

Stepwise Multiple Regression—SAS

As before, click on:

Statistics
Regression
Linear

Also as done before, highlight A, B, C, and D and move them into the
Explanatory box. Highlight F and move it into the Dependent box. Now click
the model button. Once in that Dialog box, click the model button. Once in
that Dialog box, click the selection method list from full model to stepwise
selection. At the top of the Dialog box, to the right of the tab labeled method
is another tab labeled criteria. Click on that and change the criteria to .5 to
enter and .9 to stay in the model. Click o& to return to the main Dialog box
and ok to run the regression.
If you want to save a step, in the model Dialog box, instead of stepwise,
select maximum R-square improvement. Then you won't have to change
the enter/stay criteria as before. The system will bring in the variables exactly
as in stepwise until all the independent variables are brought in and, obvi-
ously, R-square will be at its maximum. With only five independent variables,
this is a useful option. However, stick with stepwise when there are a large
number of possible independent variables.
See Table 14.13 for the SAS stepwise output.

CONCLUSION

Multiple regression is the first of several techniques for studying large


numbers of variables. For that reason, it is often excluded from introductory
Statistics and methodology texts. Yet its usage in political, social, and eco-
nomic analysis is so widespread that it is to your advantage to know about
this procedure.
This means that you are at the stage where you must depend on a
computer to do the number crunching for you. Although it is a pleasure to
let the calculations go, that pleasure is soon replaced by the pain (or at least
the challenge) of figuring out what the data mean. Why, for instance, in our
Additional Aspects of Correlation and Regression Analysis » 553

Table 14.13 Stepwise Multiple Regression—SAS


The REG Procedure
Model: MODEL 1
Dependent Variable: F

Number of Observations Read 25


Number of Observations Used ~ Zo
Stepwise Selection: Step 1
Variable B Entered: R-Square = 0.4487 and C(p) = 13.3261

Analysis of Variance

Sum of Mean
Source DF Squares Square F Value brit

Model 1 370.19829 370.19829 18.72 0.0002


Error 20 454.76171 I 7225
Corrected Total 24 824.96000

Parameter Standard
Variable Estimate Error Type Ll SS F Value JR

Intercept OO ia 2.07722 450.68274 EBay) <.0001


B 2.78177 0.64288 370.19829 18.72 0.0002

Bounds on condition number: 1, 1

The REG Procedure


Model: MODEL 1
Dependent Variable: F
Stepwise Selection: Step 2
Variable E Entered: R-Square = 0.5946 and C(p) = 6.2438

Analysis of Variance

Sum of Mean
Source IBF Squares Square F Value LP Se

Model 2 490.52312 245.26156 16.13 <.0001


Error pap) 334.43688 15.20168
Corrected Total 24 824.96000
Parameter Standard
Variable Estimate Error Type Il SS F Value Pr>F

Intercept 4.22677 2.72184 36.65924 2.41 0.1347


B 2:78591 0.56370 371.29862 24.42 <.0001
E 1.65069 0.58672 120.32482 Vee 0.0101

Bounds on condition number: 1, 4


554 <4 STATISTICS FOR THE SOCIAL SCIENCES

Table 14.13 (Continued)


Stepwise Selection: Step 3
Variable C Entered: R-Square = 0.6566 and C(p) = 4.3861

Analysis of Variance

Sum of Mean
Source DE Squares Square F Value (EPEAe

Model 3) 541.63007 180.54356 13.38 <.0001


Error 21 239.92990 13.49187
Corrected Total 24 824.96000

Parameter Standard
Variable Estimate Error Type I SS F Value Ie Rene

Intercept 1.22459 2.99241 2.25949 8 0.6805


B 2.02158 0.66049 126.39170 O57 0.0059
C L6557 7 0.95349 51.10756 oe, 0.0051
E 1.89907 0.56728 151.20044 TE Zeal 0.0031

Bounds on condition number: 1.6011, 12.604

Stepwise Selection: Step 4


Variable D Entered: R-Square = 0.0805 and C(p) = 4.5218
Analysis of Variance
Sum of Mean
Source DF Squares Square F Value Pion

Model 4 566.32995 141.58249 10.95 <.0001


Error 20 258.63005 12.93150
Corrected Total 24 824.96000

The REG Procedure


Model; MODEL 1
Dependent Variable: F
Stepwise Selection: Step 4

Parameter Standard
Variable Estimate Error Type II SS F Value Pp Se

Intercept 4.42793 3.73564 18.16850 1.40 0.2498


B 1.91110 0.65155 111.25439 8.60 0.0082
G 1.47457 0.973303 29.67673 2.2 0.1454
D =T 1556) 0.82025 24.69928 eel 0.1822
E 1.77134 0.56302 127.99987 9.90 0.0051

Bounds on condition number: 1.7408, 22.52


Additional Aspects of Correlation and Regression Analysis 555

Stepwise Selection: Step 5


Variable A Entered: R-Square = 0.6949 and C(p) = 6.0000

Analysis of Variance

? Sum of Mean
Source DF Squares Square F Value ee Se:

Model 5 573.24243 114.64849 8.05 0.0002


Error i) BS WLU ST 13.24829
Corrected Total 24 824.96000

Parameter Standard
Variable Estimate Error Type II SS F Value Pr>F

Intercept 3.55107 Oe 10.59330 0.80 0.3824


A 0.44254 0.61265 6.91248 0.52 0.4789
B 2.00881 0.67322 117.95890 8.90 0.0076
C 1.38485 0.99303 25.76604 94 OnID2
D —1.18162 0.83289 26.66457 2AM 0.1722
E 1.87974 0.58930 134.79765 10.17 0.0048

Bounds on condition number: 1.7685, 34.791


All variables left in the model are significant at the 0.9000 level.
All variables have been entered into the model.

The REG Procedure


Model: MODEL 1
Dependent Variable: F

Summary of Stepwise Selection

Variable Variable Number — Fartial Model


Entered Removed VarsIn R-Square R-Square C(p) F Value Pr>F

B 1 0.4487 0.4487 13.3261 18.72 0.0002


E 2 0.1459 0.5946 6.2438 7.92 0.0101
C 3 0.0620 0.6566 4.3861 3.79 0.0651
D 4 0.0299 0.6865 4.5218 1.91 0.1822
A 5 0.0084 0.6949 6.0000 0.52 0.4789

(Continued)
556 @ STATISTICS FOR THE SOCIAL SCIENCES

Table 14.13 (Continued)

The REG Procedure


Model; MODEL 1
Dependent Variable: F

Number of Observations Read 25


Number of Observations Used 25

Analysis of Variance

Sum of Mean
Source IDF Squares Square F Value Bret

Model 5 573.24243 114.64849 8.65 0.0002


Error 1 Zola ie 13.24829
Corrected Total 24 824.96000

Root MSE 3.63982 R-Square 0.6949


Dependent Mean 18.04000 Adj R-Sq 0.6146
Coeff Var 20.17639

Parameter Estimates

Squared Squared
Parameter Standard Partial Partial
Variable DF Estimate Error tValue Pr>|t| CorrTypel Corr Type Il

Intercept 1 SoU OTT ZA 0.89 0.3824


A 1 0.44254 0.61265 0.72 0.4789 0.01735 0.02673
B 1 2.00881 0.67322 2.95 0.0076 045921 0.31909
C 1 1.38485 0.99303 1 On 92 0.04714 0.09286
D LHS WSIG2 083289" — 142 Ony22 0.10770 0.09578
B il 1.87974 0.58930 ail) 0.0048 0.34875 0.34875

simulated data, was there no correlation to speak of between a judge’s posi-


tion on the right-to-die issue and the judge’s overall judicial ideology? Also,
why did the consumer protection factor appear not to relate to JLIB?
Sometimes it’s easier to grind Out answers and have your instructor tell you
whether they are right than it is to look at the answers on a printout and
figure out their meaning!
Beyond multiple regression are a variety of other multivariate tech-
niques with such strange names as cluster analysis, factor analysis, PROBIT,
and LISREL. And there are others on the way. Each new technique opens up
Additional Aspects of Correlation and Regression Analysis j» 557

more vistas for data analysis but also poses new problems. Some techniques
are tried and discarded. Sometimes they are resurrected later and some-
times not. Some, like multiple regression, have staying power. You are likely
to encounter these multivariate techniques again should you pursue addi-
tional coursework.
If this will be your last encounter with a statistics course, I hope you
have gained an appreciation of the fact that much of the content you have
covered is a venture into applied logic. The numbers are only the symbols
used, the language in which the logic is applied and then communicated to
others. Perhaps you will be aware that in many endeavors, you are applying
the same logical process used here, whether you are buying a car, selecting
a sofa, or listening to some expert tell you that some scientific finding is sta-
tistically significant. You can now remind that expert that there exists at least
some probability that the finding could be wrong.

Chapter 14: Summary of Major Formulas

The Conversion Between F and the Coefficient of Determination

r
> F
= —____ af =1 and n—2
N= 2k ;

Finding the Critical Value of r for Testing Statistical Significance

F critical
Veiueal = =
n — 2+ Feritical

First-Order Partial Correlation Coefficients (Three Variables: x,, x,, and x,)

AIP ANGERS GSES NMG) LBS VAC)


To — LIB —
; = 239)
5 5
f1—73; (1-73 ie ee i724 Ls

The Coefficient of Multiple Determination (Three Variables With x,


Dependent)

2 2 2 Me 2 sale
Kee = ete (li OD Ra. Tix tT Figg (l-7rj)
558 << STATISTICS FOR THE SOCIAL SCIENCES

EXERCISES
Exercise 14.1
Using Table 14.1, test the following correlations for statistical significance. Assume
nondirectional alternative hypotheses, unless told otherwise.

le re —22, fia LO
2. r=—A9, He 25
3.0 = 85, HLS
A f=—.10, n= 80
5 r=. 36, n= 100
6 f= 62, n= 45
Je ieee Be a4) one-tailed H,
6, r=—/0, few 6 directional H,
Oe pen 19, n= 102
O26, 50) fie 12 directional H,

Exercise 14.2
Calculate all partial correlations and derive the most logical causal model for each
of the following.

1. x, = Age, x, = Social Conservatism, x, = Religiosity. r,, = .93, r,, = .64, and


fh, 2.09,
2. x, =Age, x, = Pro-Choice Attitude Scale, x, = Religiosity. r,, = -.16, r,, = .90,
and r,, =—.18.
3. x, = Conservatism, x, = Pro-Life Attitude Scale, x, = Pacifism (Antiwar Attitude
Scale). r,, = .58, ,, =—.25, and r,, = .02.
4. x, = Socioeconomic Status, x, = Support for Gun Control Legislation, x, = Support
for Capital Punishment. r,, = .58, r,, =-.25, and r,, =-.43.

Exercise 14.3
Calculate and interpret the coefficients of multiple determination for predict-
ing x, from x, and x, (that is, R{,,) for each of the four data sets presented in
Exercise 14.2.

Exercise 14.4
A researcher is studying the factors that make younger voters support a given
candidate in an election. The higher one’s affinity score for a candidate,
the more likely one will be to vote for the candidate. Following are the variables
studied.
Additional Aspects of Correlation and Regression Analysis 5)5)8)

x, = Candidate Affinity (0 = low to 20 = high)


x, = Partisanship (0 = the voter totally opposes the candidate’s party to 20 =
the
voter totally supports the candidate’s party)
x, = The voter's years of formal schooling
x, = The voter's age
es

Here are summary statistics:

x,= 10.0 s, = 2.510


X,=9.9 5, = 3.131
X,= 12.5 $, = 5.252
X,= 20.5
aS
5,= 17.194
The multiple regression formula for predicting x, from the other variables was
obtained by computer.

x= 2.756 4-780x, —-.027x, — 007%,


Use the formula to predict x, when x,, x,, and x, are at their respective mean
values. How close does it come to X,?
Predict x, when x,,x,,and x, are all 0.
Generate
the betas for x,, x,,and x, and interpret their meanings.
Interpret the meaning of b, = .780 and compare it to your interpretation of
the corresponding beta.
What is the most important independent variable? What is the least important?
What other variables would you personally like to add to the model for possi-
ble explanation of candidate affinity? (Remember that these must be subject to
interval-level measurement to work in a regression.)

Exercise 14.5
An economist is trying to predict annual increases in the CPI, the Consumer Price
Index, which indicates the rate of inflation. The independent variables used are
MONEYMKT, the average annualized yields from money market funds; HOME, the
change in the average home price from year to year; and WHEAT, the year-to-year
change in wheat futures (an indicator of anticipated changes in the cost of that
commodity). A computer generates the table on the following page.
In effect, the model shows the “contribution” to the CPI of three different types of
commodities: the “cost” of money (interest), the cost of housing, and the cost of food.

q, Using only the first two decimal places, write out the regression formula for
predicting the CPI.
What will be the predicted change in the CPI in a year when money market
yields, home prices, and wheat future prices do not change?
560 << STATISTICS FOR THE SOCIAL SCIENCES

3. Predict the CPI for a year when money market yields are up 7% (plug in 7, not
.07), housing prices increase by 4%, and wheat futures decline 3%.
4. Predict the CPI for a year when money market yields decline 3%, housing
prices increase 2%, and wheat futures rise 10%.
5. Examine the betas (look at the column titled STANDARDIZED ESTIMATE) and
interpret them. What has the greatest impact on CPI? The second greatest?
The least?

Table E14.5 = The SAS System


MODEL: MODEL
DEPENDENT VARIABLE: CPI

ANALYSIS OF VARIANCE

SUM OF MEAN
SOURCE DF SQUARES SQUARE F VALUE PROB > F
MODEL 3 139.80704 46.60235 103.832 0.0001
ERROR 6 2.69296 0.44883
C TOTAL 9 142.50000
ROOT MSE 0.66994 R-SQUARE 0.9811
DEP MEAN 6.50000 ADJ R-SQ 0.9717
CV. 10.30684

PARAMETER ESTIMATES

PARAMETER STANDARD T FOR Ho: STANDARDIZED


VARIABLE DF ESTIMATE ERROR PARAMETER=0 PROB>|T| — ESTIMATE

INTERCEP 1 = -=3.839252 0.74232459 SAT 2 0.0021 0.00000000


MONEYMKT 1 0.687741 0.07243392 9.495 0.0001 0.60121698
HOME | 0.459485 0.07263228 6,326 0.0007 0.52155505
WHEAT | 0.088931 0.01873714 4.746 0.0032 0.40916409

Exercise 14.6
From the data presented in Table 14.5, pick a crime other than arson and use the
computer to find its correlation with sentence. Then pick a control variable and
find the first-order partial correlation. Try to interpret it. Add a second control
variable, run the second-order partial, and try to interpret the findings.

Exercise 14.7
From the data presented in Table 14.5, take the crime you chose as your dependent
variable in Exercise 14.6. Select three other crimes from the list as independent vari-
ables and run a stepwise multiple regression on the computer. What conclusions
can you reach?
RTE OOH RCE IEE ERED EEE IE EEA ELT ESTDEEERE LEE BEA ELLE LAE EE IEEE RES ETT EEL LERESSBOE LEELENEGOS NCL LULL NBERS AINSI ALTERS
Appendixes » 561

Appendix 1 Proportions of Area Under Standard Normal Curve

A B C A B és A B G
\ /
Zz N J) Wea a3 ;b Zz yy Th a

0.00 .0000 5000 0.30he" 1179 3821 0.60 .2257 2743


0.01 0040 4960 “Ojiuel eters 3783 0.61 .2291 2709
0.02 .0080 4920 O32 © 1255 3745 0.62 2324 2676
0.03 0120 4880 0.33 1293 3707 0.63. 2357 2643
0.04 —-.0160 4840 0.34 1331 3669 0.64 — .2389 2611
0.05 .0199 4801 0.35 .1368 3632 O60 eeedae2 2578
0.06 .0239 4761 0.36 .1406 3594 0.66 2454 2546
0.07 EY .0279 A721 0.37 1443 3557 0.67 —-.2486 2514
0.08 .0319 4681 0.38 .1480 3520 0.68718! 2517 2483
0.09 .0359 4641 (3 0eee 1517, 3483 0.69 — .2549 2451
0.10 .0398 4602 0.40 .1554 3446 0.70 .2580 2420
0.11 .0438 4562 0.41 1591 3409 C7 22611 2389
0.12 .0478 4522 0.42 1628 3372 0.72 .2642 2358
0. 15ee.0517 4483 0.43 .1664 3336 0.73. 2673 2327
0.14" (0557 4443 0.44 1700 3300 0.74 .2704 2296
0.15 .0596 4404 0.45 1736 3264 Oly5iee © 2754 2266
0.16 .0636 4364 046° 8) 172 3228 0.76 2764 2236
Oe 0675 4325 0.47. .1808 3192 0.77. .2794 2206
Giigeets 0714 4286 0.48 .1844 3156 0.78 2823 2QNT7
0.19 .0753 4247 0.49 .1879 3121 Gian We 2852 2148

0.20 0793 4207 0.50 1915 3085 0.80 .2881 2119


0.21 ~ .0832 4168 0.51 1950 3050 0.81 .2910 2090
0.220" .0871 4129 0.52. .1985 3015 0.82 .2939 2061
O25 -0910 4090 0.53. —-.2019 2981 0.83. .2967 2033
0.24 .0948 4052 0.54 2054 2946 0.84 — .2995 2005

0.25 .0987 4013 0.55 .2088 2912 0.85 3023 1977


0.26 .1026 3974 0.56 .2123 2877 0.86 3051 1949
0.27 1064 3936 Gisyaer 2157 2843 0.87. 3078 1922
0.28 1103 3897 0.58 .2190 2810 0.88 3106 1894
O20" rit 3859 Q59men 2224 2776 0.89 3133 1867

oN See Ces er eee Sevan (Continued)


562 << STATISTICS FOR THE SOCIAL SCIENCES

Appendix 1 (Continued)
G A B G

- SKY’ « SKE bh 2 es nee

0.90 Sule 1841 au 3869 Aeejil MS 4357 0643


0.91 3186 1814 122. 3888 mG, 1-53 4370 0630
0.92 soll 1788 28 OF 1093 1.54 4382 0618
0.93 3238 ANTE 1.24 ES) SOS: i>) 4394 0606
0.94 3264 1736 HES) 3944 1056 1.56 4406 0594

0.95 3289 sala 1:26 3962 .1038 iii 4418 0582


0.96 ols 1685 Aah 3980 .1020 esis) 4429 0571
Oy 3340 .1660 1.28 oT, 1003 159 444] 0559
0.98 3365 1635 1.29 4015 0985 1.60 4452 0548
0.99 3389 mlopel 1.30 4032 .0968 1.61 4463 0537

1.00 3413 allstei7/! 1.31 4049 0951 1.62 4474 0526


LO 3438 1562 Lae 4066 .0934 1.63 4484 0516
OZ 3461 SBS) 1e55 4082 0918 1.64 4495 .0505
1.03 3485 AAS 1.34 4099 0901 1.65 4505 0495
1.04 3508 .1492 oS 4115 0885 1.66 4515 0485

1.05 Dell 1469 E36 4131 0869 1.67 4525 0475


1.06 3554 1446 leaws 4147 0853 1.68 4535 0465
1.07 Soe, 1423 1.38 4162 .0838 1.69 4545 0455
1.08 OSD) 1401 Wes) Alii .0823 AG 4554 0446
1.09 3621 372 1.40 4192 .0808 (GAL 4564 0436

LELO 3643 DS7 eel .4207 .0793 1.72 4573 0427


inal 3665 BOSE) 1.42 4222 .0778 VHS 4582 0418
Wily? 3686 1314 1.43 4236 0764 1.74 4591 0409
ials) 3708 sl PAS 1.44 4251 0749 eS 4599 0401
1.14 3729 zal 1.45 42605 10735 1.76 4608 0392

IAS, 3749 AVASSII 1.46 4279 O72 len 4616 0384


1.16 KY) 30) 1.47 4292 .0708 Late 4625 0375
iy 3790 .1210 1.48 4306 .0694 1.79 4633 0367
1.18 3810 .1190 1.49 .4319 0681 1.80 4641 0359
1.19 3830 .1170 1.50 4332 .0668 i fsx 4649 0351
1.20 3849 ley sil 4345 0655 1.82 4656 0344

A B G A B C A B C
\ aes

-z A “ A \ _ —ey = Pao -z A. rS4


Appendixes ® 563

A
A

ep
B G

hae oeeee z
/
JK
ue
Y/Y.’
1.83 4004 .0336 PeeWG} 4834 0166 2.43 4925 0075
1.84 .4671 0329 2.14 .4838 0162 2.44 4927 .0073
1.85 4078 10322 "2 15 4842 0158 2.45 4929 0071
1.86 4686 .0314 2.16 4846 0154 2.46 4931 .0069
1.87 4693 .0307 Za 4850 0150 2.47 4932 .0068
1.88 4699 0301 2.18 4854 0146 2.48 4934 .0066
1.89 .4706 0294 PoeMO) 4857 0143 2.49 4936 .0064
1.90 4713 0287 2.20 4861 O39 2.50 4938 0062
198 4719 0281 Zee 4864 0136 2.51 4940 .0060
122 4726 0274 Pee 4868 W152 2.52 4941 0059
195 4732 .0268 225 4871 0129 A038 4943 0057
1.94 4738 .0262 2.24 4875 0125 2.54 A945 0055
1.95 4744 .0256 Zee) 4878 0122 4946 .0054
1.96 .4750 0250 2.20 .4881 cOMAD: 4948 0052
197) 4756 0244 22h 4884 0116 4949 0051
1.98 4761 0239 2.28 4887 0113 4951 .0049
ee) 4767 0233 BghD) .4890 0110 4952 .0048
2.00 4772 0228 2.30 4893 .0107 4953 0047
2.08 4778 0222 Zoi 4896 .0104 AD55 0045
Palys .4783 .0217 2.32 4898 .0102 4956 .0044
Z.05 4788 0212 250 .4901 0099 A957 .0043
2.04 4793 0207 2.34 4904 .0096 A959 0041
2305-4798 .0202 2.35 4906 .0094 4960 .0040
2.06 .4803 HOURSBi 2.36, 4909 0091 4961 0039
2.07 .4808 ONOZ 2.37 A911 0089 4962 .0038
2.08 4812 .0188
2.38 4913 .0087 4963 0037
2.09 4817 0183 239 .4916 .0084 4964 .0036
2.10 4821 .0179 2.40 4918 0082 4965 0035
Patil 4826 0174 2.41 4920 0080 4966 .0034
aAZ .4830 .0170 2.42 4922 .0078 4967 .0033

y ae
A B A

ei
AY we‘\
(Continued)
564 @ STATISTICS FOR THE SOCIAL SCIENCES

Appendix 1 (Continued)

A B G; A B G, A B G
/ i \
i) = / iN =, / 4 JK ST we AN

BS 4968 .0032 2.94 4984 .0016 Syl 4992 .0008


2.74 4969 0031 BSS 4984 .0016 Sal 4992 .0008
2 4970 0030 2.96 4985 0015 Bley, 4992 .0008
2.76 4971 0029 Om, 4985 0015 3.18 4993 .0007
2 is 4972 .0028 2.98 4986 .0014 Syl) 4993, .0007

IETS) 4973 .0027 299) 4986 .0014 3.20 4993 .0007


279 4974 .0026 3.00 4987 .0013 Dre 4993 .0007
2.80 4974 0026 3.01 4987 .0013 22a 4994 .0006
2.81 AOTS .0025 3.02 4987 .0013 3:25 4994 .0006
DeS2} 4976 0024 3.03 4988 .0012 3.24 4994 .0006
2.83 4977 .0023 3.04 4988 .0012 3.25 4994 .0006

2.84 4977 .0023 3.05 4989 0011 3.30 4995 .0005


Bs) 4978 .0022 3.06 4989 0011 BOS) .4996 .0004
2.86 4979 .0021 3.07 4989 .0011 3.40 4997 .0003
2.87 4979 .0021 3.08 4990 .0010 3.45 A907 .0003
2.88 4980 .0020 3.09 .4990 .0010 S50 .4998 .0002

2259) 4981 .0019 Ballo .4990 .0010 3.60 4998 .0002


2.90 4981 0019 eigilitl 4991 0009 3.70 4999 .0001
ZO 4982 .0018 Dil? 4991 .0009 3.80 4999 0001
Oe 4982 .0018 Dolls) 4991 .0009 50) 4999 0000
ES 4983 .0017 3.14 4992 .0008 4.00 4999 .0000

ACL PACT CA
A B G A A B

SOURCE: Abridged from R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricultural and
Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint of Pearson Education.
Appendixes » 565

Appendix 2 Distribution of t

Level of Significance for One-Tailed Test

10 105: O25) Ol 005 .0005

Level of Significance for Two-Tailed Test

& 20 10 . O05 02 O1 O01

3.078 6.314 12.706 31.821 63.657 636.619


1.886 2.920 4.303 6.965 9.925 31.598
1.638 2.353: in so 4.541 5.841 12.941
L539 Peay? 22776 3.747 4.604 8.610
1.476 2.015 Pas wi 3.305 4.032 6.859
1.440 1.943 2.447 3.143 52/07 5.959
1.415 13895 2.305 2.998 3.499 5.405
doo 1.860 2.306 2.896 3.395 5.041
1.383 1.833 D202 2.521 3.250 4.781
ranOrxso
Oe
me
Aan
SI
co
Ne 1.372 1.812 2228 2.764 S169 4.587
1.363 1.796 2-201 ZA 16 3.106 4,437
ao een
OR NG as 1.782 PeANTES, 2.681 3.055 4.318
1350 Naf 7Al 2.160 2.050 SOL 4.221
1.345 Lol 2.145 2.624 oat? 4.140
1.341 750 BAS 2.602 2.947 4.073
aA!
OD
NW 1537 1.746 Zale) 2.583 2.921 4.015
3350 1.740 ZolaiO 2.567 2.898 3.965
17950) 1.734 2.101 2.592 2.878 3.922
1.328 1 72o. 2.093 2.299 2.861 3.883
1.325 12> 2.086 2.528 2.845 3.850
G25 IZA 2.080 2.518 2.831 3.819
1321 eels 2.074 2.508 2.819 D192
os) 1.714 2.069 2.500 2.807 orlee
1,318 iL WALL 2.064 2.492 2197 3.745
110 1.708 2.060 2.485 Pea oh ays)
1315 1.706 2.056 2.479 arebe ja OF,
1.314 703 2.052 2.473 Pacfiiil 3.690
KF
NF ONAN S13
NONNNNNNN
KRONE
x]
OW
OO 1.701 2.048 2.467 2.763 3.674
zo ilSyl 1.699 2.045 2.462 2.756 3.659
1.310 LiGoy 2.042 2.457 2.750 3.646
1.303 1.684 2.021 2.423 2.704 Dido
17296 1.671 2.000 2.390 2.660 3.460
1.289 1.058 1.980 2.358 PASHIG ese iis)
1.282 1.645 1.960 2.326 25/6 S204

SOURCE: Abridged from Table V of R. A. Fisher and F. Yates, Statistical Tables for Biological,
an imprint
Agricultural and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley,
of Pearson Education.
566 << STATISTICS FOR THE SOCIAL SCIENCES

Appendix 3. Critical Values of F forp= .05

1 2 3 4 5} 6 8 12 24 oo

161.40 199.50 215.70 224.60 230.20 234.00 238.90 243.90 249.00 254.30

18.51 19.00 19.16 19°25 19.30 LOIDo ior 19AL 19.45 19.50

10.13 2D) 9.28 DZ 9.01 8.94 8.84 8.74 8.04 8.53

det 6.94 6.59 6.39 6.26 6.16 6.04 Sr) Silt 5.63

6.61 Sus 5.41 Sly) 5.05 4.95 4.82 4.68 ADD 4.36

Doo) 5.14 4.76 4.53 4.39 4.28 4.15 4.00 3.84 3.67

Di? 4.74 4.35 4.12 Sih 3.87 ifs) St 3.41 Re)

Doe 4.46 4.07 3.84 3.09 338: 3.44 3.28 Baz eo)

Se 4.26 3.86 BLOB) 3.48 oe 3.29 3.07 2.90 2k


4.96 4.10 anil 3.48 3190) S22, 3.07 Papell 2.74 2.54
4.84 aneks! She: 3.36 370 3.09 FESS) 29) 2.61 2.40
4.75 3.88 3.49 .26 Small 3.00 2.85 2.69 2.50 2.30
4.67 3.80 3.41 3.18 3.02 ZD2 Ziti 2.60 2.42 Zee
4.60 3.74 3.34 Hell 2.96 2.85 2.70 2259 235 21
4.54 3.08 O29) 3.06 2.90 Zany) 2.64 2.48 229) 2.07
4,49 2103 3.24 3.01 2.85 2.74 PASS) 2.42 2.24 2.01
4.45 3:52 3.20 2.96 2.81 2.70 Z.5> 2.38 Pato) 1.96
4.41 Bye) S216 2.05 Ze 2.66 Zoi Deo: Zeid 1.92
4.38 5.52 ay Al) 2.90 2.74 2.63 2.48 2.31 2 1.88
4.35 3:49 3.10 2.87 27 2.60 2.45 2.28 2.08 1.84
4.32 3.47 S107 2.84 2.08 ZY 2.42 225 2105 1.81
4.30 3.44 3.05 2.82 2.66 DiS15) 2.40 Do 2.03 1.78
4.28 3.42 3.03 2.80 2.64 250 2.38 2.20 2.00 1.76
4.26 3.40 3.01 2.78 2.62 jay 2.36 PINS: 1.98 1.73
4.24 3,98 20) 2.76 2.60 2.49 2.34 2.16 1.96 gat
4.22 Sot 2.98 2.74 ZH) 2.47 ZO OS EOS 1.69
4.21 5D 205 ZS PV 2.46 2.30 EAN) 195) 1.67
4.20 3.34 295 PoefA 2.56 2.44 2.29) Pe 1.91 1.65
4.18 DIO A312) ARKD) 2.54 2.43 2.28 2.10 1.90 1.64
4.17 Bron Uso 2.69 Zo 2.42 P| 2.09 1.89 1.62
4.08 D2) 2.84 2.61 2.45 2.34 2.18 2.00 a7 S (sill
4.00 Bulls} 2.76 2.52 PD ony DFAS) 2.10 1.92 1.70 IG Sho)
D2 3.07 2.08 2.45 229) 27 2.02 1.83 1.61 i525)
3.84 ZOD 2.60 2.37 221 2.09 Wages eS) MCS 1.00
Appendixes 567

Appendix 3 Continued—Critical Values of F for p = .01

n\n, 1 2 3 4 5 6 8 12 24 oo

1 4052 4999 5403. 5625 5764 5859 5981 6106 6234 366
2 98.49 99.01 9917 99.2 99.3 99.3 99.3 99.4 99.4 9.50
3 34.12 3081 29.46 28.71 2824 27.91 27.49 27.05 2660 26.12
4 21.20 18.00 1669 1598 1552 15.21 1480 1437 13.93 3.46
5 16.26 1327 12.06 1139 1097 10.67 10.27 9.89 9.47 9.02
6 13.74 10.92 9.78 9.15 8.75 8.47 8.10 re 7.31 6.88
7 12.25 9.55 8.45 7.85 7.46 7.19 6.84 6.47 6.07 5.65
8 11.26 8.65 7.59 7.0 6.63 6.37 6.03 5.67 5.28 4.86
9 10.56 8.02 6.99 6.42 6.06 5.80 5.47 5.11 4.73 431
10 10.04 7.56 6.55 5.99 5.64 5.39 5.06 471 4.33 3.91
(igh 9.65 7.20 6.22 5.67 5.32 5.07 4.74 4.40 4.02 3.60
12 9.33 6.93 5.95 5.41 5.06 4.82 4.50 4.16 3.78 3.36
13 9.07 6.70 5.74 5.20 4.86 4.62 4.30 3.96 3.59 3.16
14 8.86 6.51 5.56 5.03 4.69 4.46 4.14 3.80 3.43 3.00
15 8.68 6.36 5.42 4.89 4.56 4.32 4.00 3.67 3.29 2.87
16 8.53 6.23 5.29 4.77 4.44 4.20 3.89 3.55 3.18 2.75
17 8.40 6.11 5.18 4.67 4.34 4.10 3.79 3.45 3.08 2.65
18 8.28 6.01 5.09 4.58 4.25 4.01 3.71 3.37 3.00 257
19 8.18 5.93 5.01 4.50 4.17 3.94 3.63 3.30 2.92 2.49
20 8.10 5.85 4.94 4.43 4.10 3.87 3.56 3.23 2.86 2.42
21 8.02 5.78 4.87 4.37 4.04 3.81 3.51 3.17 2.80 2.36
22 7.94 5.72 4.82 431 3.99 3.76 3.45 3.12 2.75 2.31
23 7.88 5.66 4.76 4.26 3.94 3.71 3.41 3.07 2.70 2.26
2h B2 5.61 4.72 4.22 3.90 3.67 3.36 3.03 2.66 2.21
25 Val 557 4.68 4.18 3.86 3.63 3.32 2.99 2.62 Dg
26 Fae: 5.53 4.64 4.14 3.82 3.59 3.29 2.96 2.58 2.13
27 7.68 5.49 4.60 4.11 3.78 3.56 3.26 2.93 2.55 2.10
28 7.64 5.45 4.57 4.07 3.75 3.53 3.23 2.90 2.52 2.06
29 7.60 5.42 4.54 4.04 3.73 3.50 3.20 2.87 2.49 2.03
7.56 5.39 4.51 4.02 3.70 3.47 3.17 2.84 2.47 2.01
30
731 5.18 4.31 3.83 3.51 3.29 2.99 2.66 2.29 1.80
40
60 7.08 4.98 4.13 3.65 3.34 3.12 2.82 2.50 2.12 1.60
6.85 4.79 3.95 3,48 3.17 2.96 2.66 2.34 1.95 1.38
120
a 6.64 4.60 3.78 3.32 3.02 2.80 251 2.18 1.79 1.00

(Continued)
568 << STATISTICS FOR THE SOCIAL SCIENCES

Appendix 3 Continued—Critical Values of F for p = .001

n\n, 1 2 3 4 5} 6 8 12 24 oo

1 405284 500000 540379 562500 576405 585937 598144 610667 623497 636619
2 998.5 999.0 DD). 999.2 999.3 I) 3, 999.4 9904 9995) SS.
3 167.5 148.5 141.1 Sve 134.6 132.8 130.6 128.3 12539) IWASYS)
4 74.14 61.25 56.18 53.44 plea 50.53 49.00 47.41 45.77 44.05
5 47.04 36.61 33.20 yt (O) 29°75 28.84 27.64 26.42 25.14 23,18
6 sy)! 27.00 Hshs7A0) Pa O18, 20.81 20.03 19.03 IWS 16.89 I)
u Zoe 21.69 18.77 AAs, 16.21 15.52 14.63 ieewAl 12.73 11.69
8 25.42 18.49 15.83 14.39 13.49 12.86 12.04 TLS 10.30 9.34
§) 22.86 16.39 10 12.56 Wil! Alii) 10.37 ey! 8.72 7.81
10 21.04 14.91 1255 11.28 10.48 O92. 9.20 8.45 7.04 6.76
11 19.69 13.81 11.56 10.35 9.58 9.05 8.35 7.63 6.85 6.00
2 18.64 PAY) 10.80 9.63 8.89 8.38 Tei 7.00 6.25 5.42
13 17.81 Zu 10.21 Oy 8.35 7.86 Teal 6.52 5.78 4.97
14 17.14 11.78 ie) 8.62 HM 7.43 6.80 Gulls 5.41 4.60
15 16.59 11.34 9.34 8.25 ton, 7.09 6.47 5.81 5.10 4.31
16 16.12 10.97 9.00 7.94 Weil 6.81 6.19 55) 4.85 4.06
Ali? S72 10.66 8.73 7.68 7.02 6.56 5.96 SyoYs 4.63 3.85
18 5).a3) LOS? 8.49 7.46 6.81 6.35 10 Syl, 4.45 3.67
WY) 15.08 10.16 8.28 7.26 6.61 6.18 Soh 4.97 4.29 hoe
20 14.82 D5 8.10 7.10 6.46 6.02 5.44 4.82 4.15 3.38
21 14.59 77 7.94 6.95 6.32 5.88 Sol 4.70 4.03 3.26
22 14.38 9.61 7.80 6.81 6.19 5.76 5,189) 4.58 5:92 Bald
23 14.19 9.47 7.67 6.69 6.08 5.65 5.09 4.48 3.82 3.05
24 14.03 9.34 U5) G59 5.98 S155) 4.99 4.39 3.74 ZH,
25 13.88 ONLZ, 7.45 6.49 5.88 5.46 4.91 2Mehl 3.06 2.89
26 13.74 az, 726 6.41 5.80 5.38 4.83 4.24 Sa) 2.82
27 13.61 9,02 a7. 6.33 Spe) Soy 4.76 4.17 BiZ 2
28 13.50 8.93 UME) 6.25 5.66 5.24 4.09 Ani 3.46 270
29 sy oh) 8.85 v2 6.19 sy) Sys) 4.04 4.05 3.41 2.64
30 13729 8.77 7.05 6.12 S153) 5.12 4.58 4.00 3.36 2.59
40 12.61 13) 6.60 5.70 Sells) 4.73 4.21 3.04 SYIGHE eg 226
60 ie 7.70 6.17 Diol 4.76 AS, 3.87 Sheil! 2.69 1.90
120 11.38 rol Ba 4.95 4.42 4.04 Sep, 3.02 2.40 1.56
oo 10.83 6.91 5.42 4.62 4.10 3.74 Dey, 2.74 BAS 1.00

SOURCE: Abridged from Table V of R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricultural
and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint of Pearson Education.
NOTE: Values of 7, and 7, represent the degrees of freedom associated with the larger and smaller estimates
of variance, respectively.
Appendixes ®» 569

Appendix 4 Critical Values of Chi-Square

df 10 0m) OL OOL

1 eye 3.84 6.64 10.83


2 4.60 Bie, 921. 13.82
3 6.25 7.81 11.34 16.27
4 Vetke' 9.49 13.28 18.47
5 9.24 11.07 15.09 20.52
6 10.64 APR) 16.81 22.46
E 202 14.07 18.48 24.32
8 13.36 is hoy| 20.09 26r12
9 14.68 16.92 21Gy 27.88
10 15.99 18.31 20.21 29.59
al IAs: 19.68 24.72 31.26
12 18.55 21.03 26.22 32.9)
13 19.81 22.36 27.69 34.53
14 21.06 23.68 29.14 36.12
15 TDN 25.00 30.58 57.70
16 23.54 26.30 32.00 S25
17 BATT Pas} 33.41 40.79
18 25.99 28.87 34.80 42.31
19 2720 30.14 36.19 43.82
20 28.41 31.41 BeeNih 45.32
Pall 29.62 32:67 38.93 46.80
e. 30.81 33.92 40.29 48.27
23 32.01 Doky, 41.64 A973
24 33.20 36.42 42.98 51.18
25 34.38 37.65 44,31 52.62
26 35.56 38.88 45.64 54.05
i, 36.74 40.11 46.96 55.48
28 37.92 41.34 48.28 56.89
790" 39.09 42.56 49.59 58.30
30 40,26 43.77 50.89 570
40 51.80 Dae 63.69 " 73.40
50 63.17 67.50 76.15 86.66
60 74.40 79.08 88.38 99.61
70 85.53 90.53: 100.42
ee
112.32
1 2 ES ee SS SS ee eS
SOURCE: Abridged from R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricultural
and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint of Pearson
Education.
570 @ STATISTICS FOR THE SOCIAL SCIENCES

Appendix 5 Critical Values of the Correlation Coefficient

Level of Significance for One-Tailed Test

OS OZ OL OOS
Level of Significance
for Two-Tailed Test

df 10 05 02 Ol
1 988 997 9995 9999
2 900 950 980 990
3 805 878 934 959
4 729 811 882 917
5 669 754 833 874
6 622 707 789 834
7 582 666 750 798
8 549 632 716 765
9 521 602 685 735
10 497 576 658 708
11 476 553 634 684
12 458 532 612 661
13 441 514 592 641
14 426 497 574 623
15 412 482 558 606
16 400 468 542 590
17 389 456 528 575
18 378 444 516 561
19 369 433 503 549
20 360 423 492 Son
Zi 352 413 482 526
22 344 404 472 515
23 337 396 462 505
24 330 388 453 496
25 323 381 445 487
26 317 374 437 479
2 311 367 430 471
28 306 361 423 463
29 301 355 416 456
30 296 349 409 449
35 275 325 381 ~.418
40 257 304 358 393
45 243 288 338 372
50 231 273 322 354
60 oul 250 295 325
70 195 232 274 303
80 183 27 256 283
90 173 205 242 267
100 164 195 230 254
SOURCE: Abridged from Table VII of R. A. Fisher and F. Yates, Statistical Tables for Biological,
Agricultural and Medical Research (6th ed.), 1974. Reading, MA: Addison-Wesley, an imprint
of Pearson Education.
Answers to
Selected Exercises

Exercise 1.3 (There could be many other hypotheses here.)

i There is a relationship between gender and salary, such that in like


occupations, men earn more than women. Dependent variable: salary
(in appropriate national currency). Independent variable: gender (male,
female). Unit of analysis: individual. I might seek data from a source
with well-defined job categories (civil service jobs, for instance, or aca-
demic rankings in similar universities), Be careful to consider the length
of time each person has held his or her job.

There is a relationship between minority status and access to housing,


such that people of minority communities have less access than others.
Dependent variable: access to housing (high, medium, low). Independent
variable: minority status. Unit of analysis: individual. Look at statistics on
integration levels in residential areas.

There is a relationship between the-language studied by students and


the existence of minority linguistic groups in the student’s country, such
that a greater proportion of students will study the minority language
than will study other languages. Dependent variable: language selected.
Independent variable: existence of minority linguistic groups. Unit of
analysis: individual. Compare the number of U.S. students studying
Spanish to the number studying other languages.

Interest groups and PACs contribute a lower percentage of campaign


contributions where such contributions are restricted by law than they
do elsewhere. Independent: campaign contribution regulations (yes, no).

p® 571
572 4 STATISTICS FOR THE SOCIAL SCIENCES

Dependent: contributions by interest groups (PACs as a proportion of


total campaign contributions). Unit: probably country (assuming those
data are available).

There is a positive relationship between the size of a country’s armed


forces and the number of chemical and biological weapons it possesses.
Independent: size of armed forces. Dependent: number of chemical and
biological weapons. Unit: country. Look at defense statistics collected by
NATO or by private organizations.

Exercise 1.4

ie Women is not a variable; it is a category of the variable gender. Also, there


is no comparison with men being made. There is a relationship between
gender and math anxiety, such that women have higher levels of math
anxiety than do men.

Don’t state the hypothesis as a question. Age and need for social services
are positively related.

Babies have the lower birth weights, not their smoking mothers. There
is a relationship between birth weight of babies and their mother’s
smoking habit, such that babies of mothers who smoke have lower birth
weights than babies of mothers who do not smoke.

There is 7o relationship posited. All this says is that a// British parties
support national health insurance. There is a relationship between
political party and support for national health insurance, such that the
Labour Party supports a wider range of such benefits than do the other
parties.

Exercise 1.5

Countries.

Geographic area.

British counties.

The kindergarten students in Mrs. Smith’s class at Apple Valley


Elementary School registered last February 7.

Employees.
Answers to Selected Exercises 573

Exercise 2.1

i Ordinal. Class intervals are of differing sizes with the upper and lower
ones also being open-ended.
Ratio. All class intervals are closed-ended and all are equal in size.
Nominal. This would have been a Likert-type ordinal scale, except that
the sequence was broken by listing “Strongly Disagree” above “Disagree.”
To make this ordinal, reorder the categories correctly as follows:

Response ii
Strongly Agree 2)
Agree 20
Unsure 5
Disagree 1s)
Strongly Disagree 5
Total 70
Ordinal. Categories range from most to least ideal. Note that we do not
know what they meant by idealism and how it was measured. (More on
that in the next chapter.)
Ratio. All categories are equal-sized closed-ended class intervals.
Ordinal. Upper class interval is open-ended and the other class intervals
are of unequal size.
Ordinal. The categories are intended to range from most to least censor-
ship. Note that the unit of analysis here is probably country or nation-state.
Nominal. Types of races are listed [Link] ordering.

Exercise 2.2

Ratio level with individual people as the units of analysis.


Ordinal. The units of analysis are individuals, probably high school students.
Probably nominal level, because not all of Canada’s parties fit neatly into
an easily defined ideological spectrum. (There is room for disagreement
here.) Individual Canadians are the units of analysis.
Ordinal. The numbers are population rankings. Country is the unit of
analysis.
574 < STATISTICS FOR THE SOCIAL SCIENCES

There are many possible ways to operationalize these variables. The following
are thoughts or suggestions.

Exercise 3.1

1 Age. You could merely request the respondent to fill in his or her age.
But if you fear incorrect information or suspect that older subjects may
lie about their ages, consider asking them to check an appropriate class
interval, such as:
Age (Please indicate)
_____ Above 70
____ 60-69
eee ner
___ 40-49
etc,
Religion. Make sure you include all relevant faiths appropriate to your
audience, as well as categories for “other” and “none.” Also consider
whether you really want to know the subject’s religion of birth or current
religion, if they are not the same.
One could check an appropriate category, such as:
2 Jeingle
_____ Married
____ Divorced
____ Widowed
Among your subjects, there may be people cohabiting but not formally
married. How will you count them? Are you interested in marriage as a
legal status or as a sociological status? What about “marriage” between
gay men or lesbian women?

One could check the appropriate party or (in the United States) a party
plus intensity scale. Make sure there are categories for independents and
(if needed) supporters of smaller parties, such as the Greens.
Pick environmental problems that you deem important: acid rain, ozone
depletion, nuclear waste, other waste, tropical rain forests, and so on. Also,
what do you think “Attitude on environmental problems” means? Attitudes
about what exists? What should exist? What should be done about what
exists? Try to differentiate the energy issue from the pollution issue. Where
are they the same? Where do they differ?
Answers to Selected Exercises 575

Exercise 3.2

1. Look for arrests (or lack thereof) for speeches or articles expressing
politically unpopular topics. Look for societal intolerance, such as inabil-
ity to publish articles with unpopular views in the private media. How
many political parties and interest groups function in the polity? To what
extent are they governmentally controlled? Is there blacklisting or other
evidence of persecution?
2. Take several similar crimes and compare the legal penalties from one
country to another. See if you can find out for each country the average
prison term actually applied in sentencing or the actual years of incarcera-
tion by type of crime. To what crimes is the death penalty applied, if any? Is
there evidence of torture (legal or not)? Other cruel and unusual penalties?

Exercise 4.1
For the oil exporters, X = 499/7 = 71.29 and for the nonexporters, x = 264/4=
66.00. Men born in the oil-exporting countries have longer life expectancies.

Exercise 4.2
Reordering the scores from high to low:

Oil Exporters Nonexporters


Kuwait ce Jordan 71
UAE 74 Lebanon 68
Bahrain % Syria 67
Qatar ae Yemen 58
Oman 70
Saudi Arabia 69
Iraq 66

For the oil exporters, Md. Pos. = (n+ D/2 = (7 + 1)/2 = S72 = 4. ihe
fourth country in the array is Qatar with a male life expectancy Ol 72
Therefore Md. =72.
For the nonexporters, Md. Pos. = (4 + 1/2 =5/2 = 2.5. The third country in
the array is Lebanon(68) and the fourth is Syria (67). Thus Md. = (68 + 67)/2 =
67.5 = 68.
the
Since 72 is greater than 68, it is still the exporting countries that have
greater life expectancies.
576 @ STATISTICS FOR THE SOCIAL SCIENCES

Exercise 4.3

The figures for the nonexporters are unaffected. For the exporters, the
mean now becomes 43 3/6 or 72.2 years.
The median position now becomes (6 + 1)/2 = 7/2 = 3.5. This position is
shared by the third and fourth countries, Bahrain and Qatar, whose male life
expectancies are 72 and 73, respectively. Thus Md. = (72+73)/2 = 145/2 =72.5.
Summarizing:

Exporters Nonexporters
Mean Tan 66.0
Median T25 68.0

Consistent results are obtained regardless of the measure used. Male life
expectancies in the oil-exporting countries are greater than in the nonex-
porting countries.

Exercise 4.4

Oil Exporters

x= y= = cy =

16 il 16 Wit
UF & 26 10
12 1 12 8
10 1 10 ¥
) 1 9 6 a
4 1 4 5
3 4 6 4
2 2 4 2
n= fail > fx = 87

x De
x 2 eee == 75]
iy ee

Md, Pos =2+1_-2S+1_M4+1_ 12, <—


Answers to Selected Exercises j» 577

Nonexporters

OG SS
a lI ee = Ce

36 36 y)
16 16 8
13 a9 7 <—
12 i
i 4
2 o =)
4 2
“ils
RRP
RP
HE
OlR
WH y= 119

Ea
ya Ge

eS eee
barista
Ma. =13

Summarizing:

Exporters Nonexporters
Mean 7.94 iee2
Median 9.00 13.00

On the average, citizens of Arab oil-exporting countries have a greater


caloric intake (that is, they eat better) than their nonexporting cousins.

Exercise 4.5

jwae |O'

2 . VI; this would be a bimodal, symmetric distribution.

ave!
4.V

5
6
578 4 STATISTICS FOR THE SOCIAL SCIENCES

Exercise 4.6

Median
Mean
Median
Mean
Mean
Median
Mean
se
og
poi
Median

Exercise 4.13

Extroversion

50-59 | 279
AV=40"| 5 57S
30-39 | 00346789
20-29 | 0124889
10-19 | 0247
0-9 | 789

Exercise 4.14

[Link]. = 15:5
For 1st Quartile, Md. Pos. =8
For 3rd Quartile, Md. Pos. =8
Md. = 30
1st Quartile = 20
3rd Quartile = 41

Exercise 5.1

iN x= 185
ws, Md. =5
o MAD: = 1.9563
4, and 5. s*= 5300.01/9 = 588.89
6. $= 2427
Answers to Selected Exercises jp 579

Exercise 5.2

Lone 19.575
2. Md.=90

3. S°=5343.756/16 = 333.985
4. s*= 333.984 (Due to rounding in 3 above, this is slightly different.)
D285 = 18/275
6. X,,<X,,, Public employees rank it higher.

Sy, > Sy, Private employees have greater variability and publicemployees
less variability.

Exercise 5.3

x = 61.86
Md. = 68
M.D. = 29.02

AMS: 5 = 727486/7 =1039:27


a s = 32.24
ei
o's
hg

Exercise 6.3

Use the middle percentage in each cell, adding to 100% for each row. An
inverse relationship.

Exercise 6.4

A strong positive relationship.

Exercise 6.5

GDP is independent. A strong positive relationship. The higher the GDP/


Capita, the greater the number of telephone lines.

Exercise 6.6

Inconsistent but somewhat positive.


580 STATISTICS FOR THE SOCIAL SCIENCES

Exercise 7.1

1. H,: “men = women and H;: Umen 4 Lwomen


2. Hy: UFrance = MUSA and H,;: uFrance # WUSA

Exercise 7.2

1. Do a one-sample z test.
2. Two-sample ¢ test.
3. This is a table. Do chi-square.

Exercise 7.3

1. You have population data for both groups. No sampling, so no test.


2. H, says all political science majors, but should say all political science
majors in the honors program.

Exercise 7.4

1. z= -2.000
2. z2=1.500
3.2 = 2.286

Exercise 7.5

Pepi 2001
ae pes) Oil
Sp = Ulere. actually =O).
4, p< .01

Exercise 7.6

Lp = .05
2, Oe 05
Op 05
4. Not significant

Exercise 7.9

2 = 127559 = 001
Answers to Selected Exercises ®» 581

Exercise 8.1

1. 2= 1,65 Area =.0495


2. Area =.5000 + .4505 =.9505
3. Area=.0517

4. Area =.0495 + .4483 =.4978

Exercise 8.2

1. Sigma known. OK to relax normality assumption. Directional H,. z= 1.055,


Pix .05
2. Normality assumption may be relaxed, but at some risk. z = 2.48,
p01
3. Normality assumption may be relaxed with only slight risk. Sigma
unknown. Do ttest. ¢ = 2.000, df= 80 (use df= 60 on the critical values
table), p = .05. If you had the critical values at 80 degrees of freedom you
could sayp < .05.

Exercise 8.3

1. For VIO, z=-9.16, p < .001

Exercise 8.4

1. For REM, t = 1.440. df= 49. Use 40 df on the table. Not significant.

Exercise 8.5

b. 232.97, p< 01
2. z= 4.67,p < .001
3. z= -1.56, not significant.

Exercise 8.6

1. The interval ranges from 45.3 to 54.7.


3. The interval ranges from 45.5 to 54.5.
582 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 8.7

1. The interval ranges from .49 to .61.

Exercise 8.8

1. .0475
Lic ley,
3: 2062
4, 1293

Exercise 8.9

2. 336 and 56

Exercise 9.1

F=1.11. Cannot reject H,. Assume equal population variances.


t = —2.058. df= 18. Not significant.

Exercise 9.3

F = 1.74. Cannot reject H,. Assume equal population variances.


t==-2,429. af = 18. Reject H,. p < .05.

Exercise 9.5

F = 15.53. Reject H). Do not assume equal population variances.


t = —3.315. df estimated at 10. Reject H,.p < .01.

Exercise 9.7

In each case the ¢ values for equal and unequal variances are the same. But
df differs and thus the conclusions aboutp differ.

Exercise 9.9

= 1.561. df= 9. Cannot reject H,.


Answers to Selected Exercises 583

Exercise 10.1

Source SS df MS F Dp
Between 1125 - it 11.28 3.96 2205
Within 25.63 9 2.85

Not Significant. (The ¢ test in Chapter 9 had a directional H1.)

Exercise 10.3

Source SS df MS F Dp
Between PEAY 2 13.60 11.06 <.01
Within 18.41 15 iy23

Scheffé’s
fi, Ix, - x,| = Critical Value Conclusion
(ica 0.93 < 1.67 Cannot reject H.
Line pe 3.03 > 1.76 Reject Hp.
1o— fp, 2.10 > 1.82 Reject A).

Exercise 10.5

Source MS F D
Between 21,462.25 46.38 < .001
Within 462.77

Exercise 10.7

Source Ae) df MS 2)
Total 6235.36 55
Between 4093.87 <.001
Within Zo ore a

Exercise 10.9

Having only the printout, all we can say is that we may reject H,. p = .0032
(between .01 and .001).
584 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 11.1

Q =-—.680, 9 = —.382

Exercise 11.2

Either Q (gamma) or @ (phi).

Exercise 11.3

O= 818, @ = 459

Exercise 11.4

Either O (gamma) or @ (phi).


Note: the measures for Exercises 5—8 are generated on the respective
printouts in each exercise

Exercise 12.1

X° = 25.86, reject H,, p < .001


=1.14, C=.75, V=.80

Exercise 12.3

x? = 20.00, reject Fp = 00%


@ =+1.00, C=+.71, V=+1.00

Exercise 12.5

For the column variable, combine High and Medium into one category and
Low and Very Low into another. y? = 20.95.

Exercise 12.7

For the four-category variable, make it two categories by combining High


with Medium and combining Low with Very Low. For the three-category vari-
able, either combine High with Medium or Medium with Low. Either way,
two f,’s will be below 5. Chi-square is invalid. (We would have to do Fisher’s
Exact Test in this situation.)
Answers to Selected Exercises p 585

Exercise 13.1

1. r=—.962

Exercise 13.2

1. b =-.730, a = 97.66, CHEM = 97.66 —.73 DRAMA


3. Emile: 87 predicted, 93 actual Dimitri: 92 predicted, 93 actual, Alexander:
31 predicted, 29 actual

Exercise 13.4

1. ALGEBRA = 30, LITER = 71.94

ALGEBRA = 80, LITER = 26.94

2. ALGEBRA = 81, LITER = 26.04


CHEM = 75, LITER = 21.97

3. GEOM = 0, LITER = 87.34

DRAMA = 100, LITER = 85.89

Exercise 13.6

Fat e025

4, For Jack, predict LITER = 44

For Jill, predict LITER = 87

Exercise 13.7
For Exercise 10.3, 7, =.627 and E, =.596.
586 << STATISTICS FOR THE SOCIAL SCIENCES

Exercise 14.1

it Not significant.

2 ~eReject
iH. p= Ol

>. Reject ,.p< .01


A, Reject Hyp = 2025

5). Not significant.

Exercise 14.2
1. 7r,, is spurious.
2. No correlations are spurious.

Exercise 14.3

Te a ea

2: Re
= AOS

Exercise 14.4

1. x, =9:996
2. B, =.973, B, =-.057

Exercise 14.5

1. CPI =—3.83 + .68 MONEYMKT + .45 HOME + .08 WHEAT

2. Up 2.49
3. MONEYMKT: .60 the most.
WHEAT: .40 the least.
Glossary

.05 level of significance The probability level classically used by


statisticians for determining that a null hypothesis may be rejected.

Absolute value The distance or difference disregarding its sign. Here, the
distance between each value of x and the mean, regardless of whether
x is greater than the mean (a positive distance) or less than the mean
(a negative distance).

Absolute zero A zero that means a complete lack of the variable being
measured rather than some arbitrarily chosen point.
Addition rule A rule by which when outcomes are mutually exclusive,
the probability of either outcome occurring is the sum of the probabilities
of each outcome occurring.
Adjusted R-Square An estimate of R-Square in the population from
which the sample was drawn.

Alternative hypothesis or research hypothesis Statement that


indicates that in the population, the means of two or more groups differ or,
in the case of a cross-tabulation, the two variables are related.

ANOVA source table A table summarizing the results of the main steps
in the ANOVA procedure.
Antecedent variable The variable initially leading to change in the
dependent variable.
Arithmetic mean What most people learn in school as “the average.”
A measure of central tendency taking into account the distances from it
of all the scores.
Array A listing from highest to lowest (or lowest to highest).

Associated When a case falls into a particular category of one concept,


such as rain for the presence of rainfall, it also falls into a particular category
of the other, such as cloudy for sky conditions.

p 587
588 < STATISTICS FOR THE SOCIAL SCIENCES

Asymmetric measure of association A measure whose value depends


on whichever variable was dependent.
Asymptotic standard error (ASE) An alternative way of converting the
calculated measure of association to a z score.
Attributes Characteristics that do not necessarily involve amounts such as
one’s gender, eye color, or religion.
Beta coefficient or beta weight A measure that is calculated when the
relative importance of the independent variables needs to be determined.

Between-groups mean square A variance estimate based on the


between-groups sum of squares.
Between-groups sum of squares The portion of the total sum of
squares that can be accounted for by the variations of the category means
about the grand mean.
Causal modeling Procedure in which we eliminate indirect correlations
and at the same time make assumptions about causation.

Causal models Schematic diagrams showing the independent, depen-


dent, and control variables and, where appropriate, positing the flow of
causation of change in the dependent variable.
Cause When one phenomenon being studied brings about the other.
Cell Intersection of a particular row and a particular column. For example,
in Table 1.3, there are five rainy days with partly cloudy sky conditions.

Central limit theorem If repeated random samples of size 7 are drawn


from a population that is normally distributed along some variablex, having
a mean wu and a standard deviation 6, then the sampling distribution of all
theoretically possible sample means will be a normal distribution having a
mean wand a standard deviation 6/7.
Chi-square test A test of significance for two variables in a cross-tabulation.
Chi-square test for contingency A test of significance primarily used in
a cross-tabulation or contingency table.
Class interval An interval that indicates the space between two end
points.

Closed-ended A class interval that has both an upper and a lower limit.
Coefficient of alienation The proportion of variation left unexplained
by the independent variable.
Glossary 589

Coefficient of determination Indicates the proportion of variation


in the dependent variable (y) that can be explained by variation in the
independent variable (x).

Coefficient of multiple determination Coefficient that measures the


proportion of variation in the dependent variable accounted for by other
variables.

Combination The total possible samples when the order of selection is


ignored.

Computational formula A formula that generates the correct answer


but does not seek to define what the concept, such as the variance is.

Concepts Ideas.

Conceptual definition A general definition of a concept such as one


would find in a textbook or dictionary.

Concordant pairs _ Pairs of responses that are consistent with the hypothesis.

Conditional probability The probability of outcome A occurring given


that outcome B has already occurred.

Confidence interval (for means and proportions) An estimated


interval within which we are “confident”—based on sampling theory—that
the parameter we are trying to estimate will fall.

Constants Particular values of a specific linear equation that remain


* constant.

Construct validity The ability of the scale to measure variables that are
theoretically related to the variable that the scale purports to measure.

Content validity The extent to which the measure covers all the gener-
ally accepted meanings of the concept.

Contingency coefficient, Tschuprow’s 7 and Cramer’s V_ Alternative


measures to compensate for possible misinterpretations of phi.

Contingency table Table that depicts a possible relationship between


the independent variable and the dependent variable.

Contingency table, table, or cross-tabulation A way of presenting


data for purposes of testing hypotheses.

Continuous variable A variable that is not limited to a finite number of


scores.
590 << STATISTICS FOR THE SOCIAL SCIENCES

Control variable The variable whose effects on the relationship between


the other variables is eliminated by finding a partial correlation.
Correlation coefficient Measure of strength of a relationship in which
data are not grouped in tables but are individual raw scores.

Correlation-regression analysis The presentation of correlation and


regression techniques together.

Criterion validity The extent to which the measure is able to predict


some criterion external to it.
Criterion variable A substitute for the term dependent variable.
Cumulative frequency A way of accounting for all the frequencies gen-
erated up to a specific value of x. It is used to determine which value of x is
at the median position.
Curvilinear versus linear relationship A curved versus a straight-line
relationship.
Data All the information we use to verify a hypothesis.
Datum or a piece of data A single piece of information.
Deciles A number that divides scores into sets of ten and indicates for
each individual studied the number of people (or other units of analysis) his
or her score exceeds.

Deduction A process of reasoning that goes from the general to the specific.
Definitional formula A formula that not only calculates the variance but
also defines or explains the concept. In the case of the variance, the formula
defines it as the average (mean) amount of the squared deviations of the
scores from the mean.
Degrees of freedom A number that is generated to make use of a table
of critical values. An additional piece of information needed for tests where
critical values vary with the problem and may be functions of such things as
sample size.
Demographic data Background information that gives the social charac-
teristics of a subject.
Demographic variables Background information on the human
subjects studied.

Dependent or matched pairs Situation in which members of one sam-


ple are not selected independently but are instead determined by the
makeup of the other sample.
Glossary ® 591

Dependent samples ¢ test The ¢ test used when the two samples are
dependent samples.

Dependent variable The variable that is being caused or explained.


Descriptive statistics Statistics in which frequency distributions or
relationships between variables are described.

Dichotomies ‘Two-category variables.

Directional alternative hypothesis (two-tailed alternative hypothesis


or two-tailed test of significance) An alternative hypothesis that does
specify which mean will be the larger one.

Discordant pairs Pairs of responses that are inconsistent with the


hypothesis.

Dispersion The extent of clustering or spread of the scores about the


mean.
Drawing asample_ The selection of subjects to be in a sample.

Empirical questions Questions that pertain to knowing “what is.”

Expected frequencies The frequencies we would expect to find if the


null hypothesis were correct.

Experiment A test of a hypothesis under laboratory-like conditions.

For F ratio The statistic generated by the analysis of variance procedure.

F test for homogeneity of variances A test, based on the sample vari-


ances, used to determine the most appropriate ¢ test formula to use.

Face validity The extent to which the measure is subjectively viewed by


knowledgeable individuals as covering the concept.

Fallacy of affirming the consequent A principle of logic that suggests


that the only way to “prove” the alternative or research hypothesis is to
demonstrate that the null hypothesis is untrue.

First-order partial correlation Correlation that controls for only one


variable.

First-order partial regression slope A regression slope between a


dependent and an independent variable where the third variable is the
control variable.

Fisher’s Exact Test An alternative test for a two-by-two table when


chi-square is invalid.
392 @ STATISTICS FOR THE SOCIAL SCIENCES

Fractile Measurement that divides scores into smaller groups of scores of


approximately equal size.
Frequencies Headcounts or tallies indicating the number of cases in a
particular category or the total number of cases measured.

Frequency distribution A tabulation that lists the variable, its categories,


and a frequency column.

Frequency polygon Connection of dots formed by drawing a straight


line from each dot to the next dot, as x increases.

Function The case where a score on the dependent variable (vy) may be
predicted from a score on the independent variable (x). The value of y is
obtained either graphically or by an equation.

Gamma A measure designed for ordinal-by-ordinal tables that may also


be used when one of the two variables is a nominal dichotomy.

Goodman and Kruskal’s tau. A measure similar to lambda.

Goodman and Kruskal’s uncertainty coefficient Measure that is


functionally similar to Goodman and Kruskal’s lambda.

Grand mean _ The mean of all scores.

Grand total The total number of cases presented in a table. For instance,
in Table 1.1, there are 30 total days being studied.

Greek letter mu (1) Represents a population’s mean.

Greek letter sigma (6) Represents a population’s standard deviation.

Grouped interval data Grouped data that are also at the interval level of
measurement.
Grouped nominal data Data that are presented as a category of the
variable listed, and the subjects are not named but are counted (grouped)
in the category into which each subject falls.

Grouped ordinal data Data that present subjects placed into ranked
categories, ordered highest to lowest (or lowest to highest).
Histogram =Graph in which bars are created from one-half unit below
each value of x to one-half unit above that value. The faxis indicates the
frequency of each score’s occurrence.
Hypotheses Statements positing possible relationships or associations
among the phenomena being studied.
Glossary ®» 593

Independent samples The composition of one sample is in no way


matched or paired to the composition of the other sample. Thus, the two
samples reflect two separate populations.
Independent variable The variable that is doing the causing or explaining.
Index A range of scores, treated as interval or ratio level of measurement,
measuring some phenomenan.
Indirect relationship Relationship in which the correlation between
two variables is due to the presence of one or more additional variables.
Individual Data that are presented in a list of each individual subject in
the study and his or her category assignment for one or more variables.
Induction A reasoning process that goes from the specific to the general.
Inferential statistics or inductive statistics The body of knowledge
that deals with tests of significance.
Interaction effects Situations where the relationship between the
dependent variable and one of the independent variables is a function
of the levels of the other independent variable.
Interval estimation An interval of scores that is established, within
which a population’s mean (or another parameter) is likely to fall, when that
parameter is being estimated from sample data.
Interval level of measurement An interval scale in which the subject
receives a numerical score rather than a ranking and where zero is an arbi-
trarily chosen point rather than lack of what is being measured. Scores
may be below zero as well as above zero.
Intervening variable The variable through which the antecedent vari-
able brings about change in the dependent variable.
Intraclass correlation coefficient and the correlation ratio Two cor-
relation coefficients designed to measure association in an analysis of variance.
Inverse relationship Relationship in which one variable increases in
magnitude as the other variable decreases.
Inversely related A condition in which most cases cluster about the off
diagonal. A high score on one variable is associated with a low score on the
other.
Items The various components (e.g., abortion, family values, etc.) used to
generate a scale or index.
Kendall’s tau-b and Kendall’s (or Stuart’s) tau-c Measures that are
similar to gamma and are symmetric measures of association.
594 << STATISTICS FOR THE SOCIAL SCIENCES

Lambda _ Designed for a table where at least one variable is nominal and is
not a dichotomy.
Lambda symmetric The average of two lambdas.
Large effects A .80 population mean difference.
Law of large numbers A law that states that if the size of the sample,
n, is sufficiently large (no less than 30; preferably no less than 50), then
the central limit theorem will apply even if the population is not normally
distributed along variable x.
Least squares method Technique that finds the equation of the line that
best fits the points of a diagram.
Levels of measurement or scales Measurement that falls into one of
four general categories: nominal, ordinal, interval, or ratio.

Likert scale A scale whose categories are based on the level of agreement
with a particular statement or issue.
Linear equation Equation in which the points generated will graph as a
straight line rather than any other kind of graphic figure.
Linear regression Technique that finds a line that “fits” the scatter
of data points in such a way as to provide for any given value of x the best
estimate of the corresponding value of y.
Linearly related Relationship that is shown as an exact straight line.
Main diagonal Diagonal line from upper left to lower right.
Marginal totals Row and column totals found in the margins of tables.
Mean deviation An average distance that a score deviates from the mean.
Mean square The mean squared deviation of a score (a value of x) from
the mean of all scores.

Measurement A very specific process, such as measuring length, but also


other simpler actions such as assignment of a person to a particular category
of a variable.
Measures of association Index numbers that generally range in mag-
nitude from 0 (measuring no association) to 1 (usually meaning perfect
association).

Measures of central tendency Averages or measures of location that


find a single number that reflects the middle of the distribution of scores—
the “average” score for that group.
Glossary 595

Measures of dispersion Measures concerning the degree that the scores


under study are dispersed or spread around the mean.
Median A value in which there are as many scores greater than the
median as there are scores less than the median.
Median position The person, place, or thing that possesses the median
score or middle position. —.
Medium effects Hypothesized .50 difference between the population
means.
Microsoft Excel Microsoft’s spreadsheet program that also may be used
for statistical analysis.
Microsoft Windows Microsoft’s widely used operating system.
Mnemonics Short names that replace algebraic letters in identifying the
variables.
Modal class or modal category Where data have been grouped, a class
interval or category that contains more cases than can be found in either
category adjacent to it.
Modality The number of modes found in the frequency distribution.
Mode acategory ofa variable that contains more cases than can be found
in either category adjacent to it.
Multiple correlation Coefficient that measures the correlation between
a dependent variable and the combined effect of other designated variables
in the system.
Multiple regression or multiple linear regression Technique of
developing predictive equations when there is more than one independent
variable present.
Multiplication rule A rule that is used to find P(A and D).
n category A term indicating more than two categories.
n-dimensional hyperplane A structure with more than three dimen-
sions that can be described mathematically.
Necessary condition A condition that must be present in order for some
outcome to occur.
Nominal Pertains to the act of naming.
Nondirectional alternative hypothesis An alternative hypothesis that
does not specify the directionality (i.e., which mean will ultimately be the
larger).
596 @ STATISTICS FOR THE SOCIAL SCIENCES

Normal distributions <A family of frequency distributions that, when


graphed, often resemble bells.

Normality assumption The assumption that that the population being


studied is normally distributed along variable x.

Normative questions Questions that pertain to “what ought to be.”

Null hypothesis A statement postulating that in the population, the


means of two or more groups are the same or, in the case of two vari-
ables in a cross-tabulation, that in the population, the two variables are
unrelated.

Observed frequencies The actual frequencies observed in the contin-


gency table.

Off diagonal Clustering on a diagonal line that goes from the upper right-
hand side of the table to the lower left-hand side.

One-sample ¢ test A test that can be performed when we know the


population’s mean but not its standard deviation.

One-sample tests Tests that compare data from a sample to similar data
in a population,

One-sample z test A test of significance that can be performed when we


know the population’s standard deviation as well as its mean.

One-way analysis of variance (one-way ANOVA) A test in which two


or more sample means may be compared simultaneously.

Open-ended A class interval that has a lower limit but no upper limit or
vice versa.

Ordered pair A set of two numbers in parentheses separated by a


comma, indicating a point on a graph.

Ordered triplet Each point in space that is referenced to the three axes.

Ordinal Involving a rank order or other ordering.

Origin The point of a graph where the two axes intersect, indicating a
value of zero on each axis.

Paired difference ¢ test ©A one-sample ¢ test applied to the differences in


each pair of scores.

Partial correlation The amount of relationship not attributable to other


variables in the system.
Glossary p 597

Partial regression slope Slope in a multiple regression.


Participants or subjects The people being studied in an experiment.
Pearson’s C and Cramer’s V_ Measures that are similar to the phi coeffi-
cient but are more accurate when applied to tables larger than 2 x 2.
Pearsonian product-moment correlation coefficient or Pearson’s r
- Coefficient that is used whemboth variables are an interval or a ratio level of
measurement.
Percentiles A number that divides scores into sets of 100 and indicates
for each individual studied the number of people (or other units of analysis)
his or her score exceeds.

Permutation The total possible samples that can be drawn from a popu-
lation when the order of selection is a factor.
Phi An alternative measure of association for a two-by-two table that is
sometimes preferable to O.
Pooled estimate of common variance Estimate based on a weighted
average of two sample variances being used to estimate the population
variance in finding the standard error.
Population or sampling universe The group about which we want to
generalize.
Population parameters Information computed from population data.
Positive relationship A relationship in which greater is associated with
greater; less with less.
Post hoc A follow-on procedure that is used once a null hypothesis has
been rejected.
Post hoc tests of multiple comparisons Tests that enable us to narrow
our conclusion to specifically where these population inequalities are to be
found.
Predictions of y The prediction or estimate of y from a specific value of
x as generated by the regression formula.
Predictor variable A substitute for the term independent variable.
Probabilities Proportions that reflect the likelihood of a particular outcome
occurring.
Proportionate reduction in error (PRE) The reduction in assignment
errors when we know all the frequencies in the table, rather than just the
totals, expressed as a proportion of the errors made when knowing just the
totals.
598 @ STATISTICS FOR THE SOCIAL SCIENCES

Qualitative or categorical data Data that are assigned to categories that


do not imply amounts.
Quantitative data Data that are assigned to categories that are involved
with amounts.
Quartiles A number that divides scores into sets of four and indicates for
each individual studied the number of people (or other units of analysis) his
or her score exceeds.
Random sample A sample that is not biased.
Range The simplest measure of dispersion that compares the highest
score and the lowest score achieved for a given set of scores.
Ratio level A level of measurement similar to interval level, but where
zero is an absolute zero, meaning none of what is being measured. Scores
may not be below zero.
Raw score A simple numerical score.
Regression equation The mechanism for estimating a y score from the
respective x score.
Regression of yon x Technique that informs us that the line we gener-
ate will be the line that enables us to most accurately predicty from x.
Reliability The likelihood that the scale is actually measuring what it is
supposed to measure.
Robust Accurate even when underlying assumptions (such as equal
population variances) are violated.
Sample The smaller group from the population that is selected to be
studied.
Sample statistics Information computed from sample data.
Sampling bias or biased sample The mechanism for selecting the sample
that causes the sample to be not representative of the population as a whole.
Sampling distribution of sample means The frequency distribution
that would be obtained from calculating the means of all theoretically
possible samples of a designated size that could be drawn from a given
population.
Sampling error A deviation from what actually exists in the population
not associated with sampling bias but still existing, even though the sample
was randomly drawn.
SAS A set of statistical computer routines and a programming language:
Statistical Analysis System.
Glossary 599

Scatter diagram (or scattergram or scatter plot) Diagram with data


points that are scattered on the graph and not a perfect line or other figure.

Scheffeé’s critical value The value in Scheffé’s test needed to reject the
null hypothesis.
Scheffe’s test A test that finds the critical difference between any
two sample means that is necessary to reject the null hypothesis that their
corresponding population means are equal.
Scientific laws Hypotheses verified so often that they have a high
probability of being correct.

Scientific method A series of logical steps that, if followed, help minimize


any distortion of facts stemming from the researcher’s personal values and
beliefs.

Scientists People who engage in collecting and interpreting empirical


information.

Scores Numbers that are used to represent amounts or rankings.

Second-order partial correlation Correlation that controls for two


variables.

Sigma-hat (6) An estimate of sigma.

Simple random sample A sample drawn in such a way that every


member of the population has an equal likelihood of being included in the
sample.

Skewed to the left or negatively skewed Skewness runs in the


direction of increasing negative values on the x-axis.

Skewed to the right or positively skewed Skewness is in the direction


of increasing positive values on the x-axis.

Skewness_ The extent to which the frequency distribution deviates from


symmetry.
Slope The change in y per unit change in x.
Small effects Hypothesized .20 difference between u for the experimental
group and pu for the control groups.
Smooth curve Connection of dots similar to a frequency polygon but
generated by a curved line fitting through the dots instead of a series of
straight lines.
Social sciences The empirical study of social phenomena.
600 << STATISTICS FOR THE SOCIAL SCIENCES

Somer’s d_ Measure that is similar to gamma but is asymmetric.

Split-half reliability A measure of internal consistency that splits an


overall scale into two scales, each containing half the original items.

SPSS A set of statistical computer routines: Statistical Package for the


Social Sciences.
Spurious Indirect relationship in which the control variable is independent.

Spurious relationship Relationship between two variables that is the


product of a common independent variable.

Standard deviation The positive square root of the variance, which


provides a measure of dispersion closer in size to the mean deviation.

Standard error of the mean or the standard error The standard


deviation of the sampling distribution, designated with the symbol o;.
Standard score A score universally designated by the letter z, in which
that score is expressed in standard deviation units from the mean.

Standardized partial regression slope A slope expressed not in the


original units used but in standard deviation units of the dependent variable.

Statistical power The likelihood that our test will reject the null hypoth-
esis when, in fact, H, really is true.
Statistics The study of how we describe and make inferences from data.

Stem and leaf display A graphic representation that combines the


visual effect of a histogram but preserves the actual scores in small- to
medium-sized data sets.

Stepwise multiple regression A procedure in which each independent


variable is added to the regression in a separate step. The order of entry is
based on criteria selected by the researcher.

Sufficient condition A condition in which the predicted outcome will


definitely take place.

Sum of squares The sum of the squared deviations of the values of x


from the mean.

Summation symbol Symbol represented by the uppercase Greek letter


sigma ()°).
Symmetric frequency distribution A frequency distribution with no
skewness (skewness equals zero).
Glossary »» 601

Symmetric measure of association A measure whose value does not


depend on whether the variable was dependent or independent.
Symmetry The balance between the right and left halves of the curve.
i test A test of significance similar to the z test but used when the popu-
lation’s standard deviation is unknown.
Temporal sequence When one phenomenon being studied occurs
earlier in time than the other.
Test-retest reliability A measure that determines an individual’s consis-
tency in responding the same way to a specific item over time.
Tests of statistical significance Techniques that help us to generalize to
a larger group.

Theory A set of interrelated hypotheses that together explain some phe-


nomenon such as why one identifies with a particular political party.
Total sum of squares The total of the squared deviations of scores about
the grand mean.

Two-sample ¢ test A? test that compares two sample means, rather than
one sample’s mean to another population’s mean.
Two-way analysis of variance Analysis of variance that includes a sec-
ond independent variable.
Type I error or alpha error The probability of falsely rejecting a true
null hypothesis.
Type IJ error or beta error The probability that the null hypothesis is
really false—H, is true—but our obtained statistic—z, 4, and so on—was too
low to enable us to reject the H,, even though it “ought to be” rejected.
Ungrouped frequency distribution Scores listed in a sequence (usually
highest to lowest) that includes every score that actually appears in our results.
Unimodal, bimodal, and trimodal A distribution with one, two, and
three modes, respectively.
Unit of analysis What we actually measure or study to test our hypothesis:
from whom or from what the measurement is made.
Validity The extent to which the concept one wishes to measure is actually
being measured by a particular scale or index.
Variables Particular values of a specific linear equation that vary from
person to person (or unit of analysis to unit of analysis).
602 < STATISTICS FOR THE SOCIAL SCIENCES

Variance An “average” or mean value of the squared deviations of the


scores from the mean.
Within-groups degrees of freedom That portion of the total degrees of
freedom not accounted for by the number of groups studied.
Within-groups mean square A population variance estimate based on
what the categories of the independent variable do not explain—the varia-
tion of scores within the groups.

Within-groups sum of squares or error sum of squares The portion


of the total sum of squares left unexplained by the variations of the category
means about the grand mean.

Working or operational definition A definition of the way that some-


one or something will be measured to determine the subject’s score on a
variable.

x-axis The axis that extends horizontally.

x-axis and f-axis Two perpendicular lines, a horizontal line labeled x and
a vertical line labeled f

x-bar (x) The mean value of the variable x.


x-coordinate The first number in an ordered pair.

Yates’s correction for continuity An adjustment of the chi-square fora


1 df table by applying a factor that lowers the value of chi-square obtained,
making it harder to reject the null hypothesis.

y-axis The axis that extends vertically.

y-coordinate The second number in an ordered pair.


y-intercept The value ofy at that point where the line crosses the y-axis
(Le where =10),

Yule’s Q The proportionate excess of concordant over discordant pairs of


observations.
z test for proportions A z test designed to test whether the difference
between proportions in a sample reflects the difference in the population.
Zero-order correlation Correlation that does not control for any other
variables.

Zero-order regression slope A Simple two-variable regression slope.


Index

Absolute deviation Asymmetric distribution, 107, 109


See Mean deviation (V.D.) Asymmetric measure of association, 372
Absolute value, 131-132 Asymptotic standard error (ASE), 426
Absolute zero, 46-47 Attributes, 36, 66-67
Addition rule, 258-260 Averages. See Central tendency;
Adjusted R-SQUARE (R?), 530-533 Mean deviation (M.D.)
Alpha error
See Type I error Behaviors, 66-67
Alternative hypotheses, 197, 210-214, Beta (f) coefficient, 528-530
234, 281-283, 422-425 Beta error. See Type II error
See also Significance tests Between-groups degrees of freedom (df,),
Analysis of variance (ANOVA) 529-330
applications, 318-322, 338-340 Between-groups mean square (MS,), 329-330
computer programs, 343-348 Between-groups sum of squares (SS,,), 328
correlation-regression analysis, 486-489, Biased sampling, 194-195
498-506 Bimodal distribution, 103, 108, 111
definition, 318 Bonferroni test, 343
F € ratio), 322-326, 337-338, Boxplots (box and whisker plots), 113-116
498-506, 557
interaction effects, 351 Cartesian coordinates, 447-451
major formulas, 352-353 Caseload analysis, 26-28
one-way analysis of variance (One-way Categorical data
ANOVA), 202, 319, 334-335 See Qualitative (categorical) data
procedures, 330-337 Causal models, 22-25, 172-174, 508-515
source table, 337 Cells (of tables), 8
terminology, 326-330 Central limit theorem, 238-244
two-way analysis of variance (two-way Central tendency
ANOVA), 348-352 applications, 98
See also Multiple regression definition, 84
Antecedent variables, 173-174 . graphic representations, 104-116
Arithmetic mean. See Mean levels of measurement, 105-106
Array, 90-92 measurement techniques, 85-103
Associations Chi-square test
association matrices, 381-385 calculation procedures, 412-415
causal models, 22-25 computer programs, 431-435
chi-square test, 425-430, 437 contingency context, 398—405
curvilinearity, 377-380, 460-462 contingency flowchart, 419
decision flowchart, 382 critical values, 408-411
definition, 9 definition, 198
major formulas, 385-386 expected frequencies, 399-408, 417-420, 436
measures of association, 360 major formulas, 436-437
n-by-n tables, 367-377, 386 measures of association, 425-430, 437
phi (@) coefficient, 365-366, 429-430 observed frequencies, 399-408
two-by-two tables, 360-366, 385-386 validity, 417-421
Yule’s QO, 362-365 Yates’s correction for continuity, 415-417

» 603
604 << STATISTICS FOR THE SOCIAL SCIENCES

Class intervals, 51-54, 99-100, 112-113, 153 multiple regression, 520-561


Closed-ended class intervals, 51-52 partial correlations, 508-515, 538-541
Clustering, 54, 128, 159, 164 predictions of y, 467-468
Coefficient of alienation (1 — 77), 474 rand b, 498-506
Coefficient of determination (7°), 446, 472-474, regression equation, 446-447, 474-479
498-506, 557 regression of x on y, 466
Coefficient of multiple determination (R’), regression of y on x, 463-468, 489
Si, Seyi stepwise multiple regression, 533-538,
Combinations, 261—262 549-556
Computational formulas, 136-140, 216 See also Pearson’s r (Pearsonian product-
Computer applications moment correlation coefficient)
analysis of variance (ANOVA), 343-348 Cramer’s V 380-381, 430
chi-square test, 431-435 Criterion validity, 73-74
correlation-regression analysis, 479-486, Criterion variables, 25
530-552 Critical values
measures of association, 380 chi-square test, 408-411
multiple regression, 530-556 F (F ratio), 331-333
partial correlations, 538-541 F test for homogeneity of variance,
statistical output programs, 175-180 277, 279-280
stepwise multiple regression, 549-556 Pearson’s r (Pearsonian product-moment
t tests, 276, 288-303 correlation coefficient), 506-508, 557
zero-order correlations, 542-543 Scheffé’s critical value, 341-342
Concepts, 9 t statistic, 248-252
Conceptual definitions, 67-68 Z Statistic, 200-207, 235, 236
Concordant pairs, 363, 367 Cross-tabulation, 7-8, 150-151
Concurrent validity Cross-tabulaton
See Criterion validity See also Computer applications
Conditional probability, 258 Cumulative frequency (cf), 95-97
Confidence intervals, 254-257, 288 Curves, normal
Constants, 455-456 See Normal distributions
Construct validity, 74 Curves, smooth, 101-102
Content validity, 73 Curvilinearity, 377-380, 460-462
Contingency chi-square test
See Chi-square test Data (datum)
Contingency coefficient definition, 14
See Pearson’s C (contingency coefficient) demographic variables, 47-48, 64-65
Contingency tables, 7-8, 150-158 grouped data, 94-97, 151-155
See also Associations interval-level data, 48-54, 99-100
Continuous variables, 102-103 nominal-level data, 38-40
Control groups, 165, 274-275 ordinal-level data, 40-42
Control variables, 164-174, 511-515 Deciles, 114
Correlation matrix Decision-making process, 206-209
See Associations Deduction, 11-12
Correlation-regression analysis Definitional formulas, 136-140, 216 .
analysis of variance (ANOVA), 486-489, Degrees of freedom (df)
498-506 between-groups degrees of freedom (df,),
applications, 444-447 329-350
Cartesian coordinates, 447-451 definition, 214
causal models, 508-515 F test for homogeneity of variance,
computer programs, 479-486, 530-552 279, 281, 284-286
correlation ratio (N), 488-489, 490 t tests, 249-250, 287-289
intraclass correlation coefficient (7), within-groups degrees of freedom (dfy),
486-488, 490 529-330
least squares method, 463-468, 489 Demographic variables, 47-48, 64-65
linearity, 451-479 Dependent samples, 274-275, 297-303, 310
major formulas, 489-490, 557 Dependent variables, 24-25, 172-174
multiple correlation (R), 516-520 See also Contingency tables
Index & 605

Descartes, René, 447 grouped data, 153-155


Descriptive statistics, 193 median (Md.), 94-97
Diagonals normal distributions, 226-234, 263
clustering, 54, 159, 164 skewness, 106-111, 115-116
main diagonals, 19, 159 standard deviation (6), 139-142, 199, 216,
off diagonals, 20-21, 159, 164 228-229, 239-241
Dichotomies, 39-40, 44 variance, 139-142
Directional alternative hypothesis, 210-214, Frequency polygons, 101-102
234, 281-283, 422-425 ‘ F test for homogeneity of variance, 276-297,
See also Significance tests 309
Discordant pairs, 363, 367 Functions, 451-452
Dispersion
definition, 84, 128 Gamma
measurement techniques, 128-141 See Goodman and Kruskal’s gamma (Y)
Gastil, Raymond, 151
Effect sizes, 308-309 Goodman and Kruskal’s gamma (y¥), 367-371
Empirical questions, 2-5 Goodman and Kruskal’s lambda (A), 371-377
Equal population variance ¢ test, 276-277, 279, Goodman and Kruskal’s uncertainty
281, 283, 286-290, 310 coefficient, 381
Error Grand mean, 323-326
asymptotic standard error (ASE), 426 Grand totals, 7-8
proportionate reduction in error (PRE), Graphs, 104-105, 447-451
374-375 Grouped data, 94-97, 151-155
sampling error, 196-197 Grouped interval data, 50-51
standard error, 239-241, 275-276 Grouped nominal data, 39
Type I error, 207—208, 307 Grouped ordinal data, 41-42
Type II error, 223, 307-308
Error sum of squares Histograms, 101-102
See Within-groups sum of squares (SS\)) Hypothesis/hypotheses
Excel. See Microsoft Excel alternative hypotheses, 197, 210-214, 234,
Expected frequencies, 399-408, 417-420, 436 281-283
Experiments, 11-12 definition, 5-6, 10-11
directional alternative hypothesis, 210-214,
Face validity, 73 234, 281-283, 422-425
Fallacy of affirming the consequent, 197-198 null hypothesis (H,), 197, 200-201, 205-209,
fraxis, 100-101 212-213, 223
F (F ratio) testing procedures, 11-15
compared to ¢ tests, 337-338 theory development, 16-18
correlation-regression analysis, 498-506, 557 See also Chi-square test; Significance tests;
critical values, 331-333 t tests
definition, 319
intuitive approach, 322-326 Independent samples, 272-286, 290-297, 309
Pearson’s 7 (Pearsonian product-moment Independent variables, 24-25, 172-174
correlation coefficient), 506-508 See also Contingency tables
First-order partial correlations, 512-515, 557 Index and scale construction, 69-73
First-order partial regression slopes, 521-522 Indirect relationships, 166, 509
Fisher’s Exact Test, 418-421 Individual data, 38, 40-41, 48-49
Fractiles, 113-114 Induction, 11
Frequencies Inductive statistics, 193
definition, 45 Inferential statistics, 192-195
expected frequencies, 399-408, 417-420, Interaction effects, 351
436 Interval estimation, 254-257
observed frequencies, 399-408 Interval level of measurement, 45-54, 73
Frequency distribution (/) See also Correlation-regression analysis
cumulative frequency (cf), 95-97 Intervening variables, 173-174, 510-511
definition, 39 Intraclass correlation coefficient (7),
graphs, 104-105 486-488, 490
606 @ STATISTICS FOR THE SOCIAL SCIENCES

Inverse relationships, 20-22, 159-164, 459-460 Medium effects, 308


Items of a scale or index, 70-71 Microsoft Excel
analysis of variance (ANOVA), 348, 349
Kendall's (or Stuart’s) tau-c, 381 correlation-regression analysis, 485-486, 487
Kendall’s tau-b, 381 definition, 175
multiple regression, 547, 548
Lambda t tests, 294-297, 303, 305
See Goodman and Kruskal’s lambda (A) Microsoft Windows, 175
Lambda symmetric Mnemonics, 525
See Goodman and Kruskal’s lambda (A) Modal class, 99-100
Large effects, 308 Modality, 103
Law of large numbers, 244-246 Mode, 98-103
Least squares method, 463-468, 489 Multiple comparisons, post hoc tests of,
Levels of measurement, 35-55 341-343
Levene’s test for equality of variances, Multiple correlation (R), 516-520
288-289 Multiple regression, 520-561
Likert scales, 43-45 Multiplication rule, 260-261
Linearity
curvilinearity, 379-380, 460-462 n-by-n tables, 367-377, 386
linear equations, 455—460 n-category nominal scale, 40
linear regression, 460-479 n-dimensional hyperplane, 519
linear regression equation, 444 Necessary conditions, 15
linear relationships, 159, 379-380, 451-455 Negative relationships
multiple regression, 520-561 See Inverse relationships
Location Nominal level of measurement, 36—40, 44,
See Central tendency 54-55
Nondirectional alternative hypothesis, 210
Main diagonal, 19, 159 Normal distributions, 226-234, 263
Marginal totals, 7-8 Normality assumption, 244-246
Matched pairs, 274-275 Normative questions, 2—5
Matrix, association Null hypothesis (H,), 197, 200-201, 205-209,
See Associations 212-213, 223
Mean See also Chi-square test; Significance tests;
comparison tests, 198-203 t tests
definition, 85-90
grand mean, 323-326 Observed frequencies, 399-408
major formulas, 117, 216 Off diagonal, 20-21, 159, 164
sampling distribution of sample means, One-sample tests, 201, 203-209
235-238 One-sample ¢ tests
Mean deviation (M.D.), 130-132, 141 See t tests
Mean square (MS), 327, 499-506 One-tailed tests
Measurement See Directional alternative hypothesis
absolute zero, 46-47 One-way analysis of variance (one-way ANOVA),
definition, 34 202, 319; 334-335
interval level of measurement, 45-54, 73 Open-ended class intervals, 52-54
levels of measurement, 35—55, 105-106 Operational definitions, 65-69, 77-79
nominal level of measurement, 36-40, 44, Ordered pairs, 449-450
54-55 Ordered triplet, 516-518
ordinal level of measurement, 40-42, 44 Ordinal level of measurement, 40-42, 44
ratio level of measurement, 46—47 Origin (ofagraph), 100, 447
See also Associations; Central tendency; Outliers, 115-116
Correlation-regression analysis
Median (Md.) Paired difference ¢ tests, 297-301
array, 90-92 Partial correlations, 508-515, 538-541
frequency distribution (f), 94-97 Partial regression slopes, 521-523, 528-530
grouped data, 94-97 Partial tables, 167-172
median position (Md. Pos.), 90-97, 117 Participants (in an experiment), 320-322
Index B® 607

Pearson’s C (contingency coefficient), Ratio level of measurement, 46-47


380-381, 430 See also Correlation-regression analysis
Pearson’s 7 (Pearsonian product-moment Raw scores, 48-49
correlation coefficient) Reasoning process
calculation procedures, 468-472 See Scientific method
critical values, 506-510 Regression analysis
definition, 444, 446 See Correlation-regression analysis
major formulas, 489, 557 Regression equation, 446-447, 474-479
measures of association, 381 ~ Relationships
Percentage generation, 155-158 curvilinearity, 377-380, 460-462
Percentiles, 114, 116 indirect relationships, 166, 509
Permutations, 261—262 inverse relationships, 20-22, 159-164,
Phi (@) coefficient, 365-366, 429-430 459-460
.05 level of significance, 205-208, 211-214 linear relationships, 159, 379-380, 451-479
Pooled estimate of common variance, 276 measures of association, 359-395
Population (sampling universe) positive relationships, 19, 21-22, 159-164,
central limit theorem, 238-244 459-460
combinations, 261—262 spurious relationships, 166, 509
definition, 193-194 variables, 18-22, 150-174
mean (1), 198-203, 216, 235-238 See also Correlation-regression analysis
normality assumption, 244-246 Relevance, 304, 306
parameters, 198-200, 216, 272, 274 Reliability, 75-77
participants, 320-322 Research hypotheses
permutations, 261-262 See Alternative hypotheses
standard deviation (6), 199, 216, 228-229, Research significance, 306
239-241 Robustness, 338
variance (07), 199, 216
Positive relationships, 19, 21-22, 159-164, Sample/sampling
459-460 bias, 194-195
Post hoc procedures, 318, 340-343, 346, 348 central limit theorem, 238-244
Predictions of y, 467-468 combinations, 261—262
Predictive validity comparison tests, 200-203
See Criterion validity definition, 194
Predictor variables, 25 dependent samples, 274-275, 297-303, 310
Probabilities error, 196-197
addition rule, 258-260 independent samples, 272-274,
conditional probability, 258 290-297, 309
critical values, 206-207, 235, 236 permutations, 261-262
definition, 206 random samples, 195-196, 238, 272-274
F test for homogeneity of variance, sampling distribution of sample means,
289-290 235-246
multiplication rule, 260-261 size effects, 308-309
0S level of significance, 205-208, 211-214 standard error, 239-241
Proportionate reduction in error (PRE), statistics, 198-200, 216
374-375 Sampling universe
Proportions See Population (sampling universe)
confidence intervals, 256-257 SAS (computer program)
z tests, 253-254, 263 analysis of variance (ANOVA), 345-348
chi-square test, 433-435
Qualitative (categorical) data, 35-36 contingency tables, 177-180
Quantitative data, 35-36 correlation-regression analysis, 483—485
Quartiles, 114-115 definition, 175
measures of association, 380
ry. See Pearson’s r (Pearsonian product-moment multiple regression, 530-533, 543, 545-546
correlation coefficient) stepwise multiple regression, 533-538,
Random samples, 195-196; 238, 272-274 552, 553-556
Range, 129-130, 153 t tests, 291-294, 301, 303, 304
608 << STATISTICS FOR THE SOCIAL SCIENCES

Scale construction, 69-73 t tests, 288-291, 301, 302


Scales, 35 zero-order correlations, 542-543
See also Measurement Spurious relationships, 166, 509
Scatter diagram, 462-463 Standard deviation
Scheffé’s critical value, 341-342 computational formulas, 136-139
Scheffé’s test, 341-343, 346, 348, 353 definition, 133-136
Scientific laws, 12 frequency distributions, 139-142
Scientific method, 2, 5—7, 29 population’s mean, 199
Scientists, 5 sampling distribution of sample means,
Scores 239-241
definition, 45 significance tests, 263
grand mean, 323-326 standard score, 228-229
grouped data, 153-154 t tests, 246-249, 275-276
index and scale construction, 69-73 Standard error, 239, 275-276
interval estimation, 254-257 Standardized partial regression slope, 528-530
raw scores, 48-49 Standard score, 228-229
standard score, 228-229 Statistical power, 306-309
Second-order partial correlations, 512 Statistical significance
Second-order partial regression slopes, 521 chi-square test, 426-429
Sigma-hat (6), 247, 263 compared to research significance,
Sigma-hat squared (G2), 286-287 303-304, 306
Significance tests limitations, 435-436
analysis of variance (ANOVA), 326-330 measures of association, 426-429
chi-square test, 426-429 Pearson’s r (Pearsonian product-moment
comparison tests, 200-203, 273-274 correlation coefficient), 506-510, 557
definition, 193 r and b, 498-506
degrees of freedom (df), 214, 249-250, 279, z tests, 216, 234-246
281, 284-289 Statistics, definition of, 29
interval estimation, 254-257 Stem and leaf displays, 112-113
limitations, 435-436 Stepwise multiple regression, 533-538,
major formulas, 263 549-556
one-sample tests, 201, 203-209, 216, 246-249 Stuart’s tau-c, 381
one-way analysis of variance (one-way Subjects (in an experiment), 320-322
ANOVA), 202, 319, 334-335 Sufficient conditions, 15
procedures, 214-215 Summation symbol (2), 86-90
two-sample ¢ tests, 202 See also Cumulative frequency (cf );
See also Analysis of variance (ANOVA); Definitional formulas
Chi-square test; Error; ¢ tests; z tests Sum of squares (SS), 327, 499-506
Simple random samples, 195-196 Symmetric distribution, 106-107, 111
Skewness, 106-111, 115-116 Symmetric measure of association, 372
Slope, 456-460, 521-524, 528-530 Symmetry, 106-107
Small effects, 308
Smooth curves, 101-102 Tables
Social sciences, definition of, 5 construction process, 151-155
Somer’s d, 381 control variables, 164-174
Split-half reliability, 76 definition, 7-8
SPSS (computer program) interpretation, 159-164
analysis of variance (ANOVA), 343-346 nominal level of measurement, 54—55
chi-square test, 431-433 partial tables, 167-172
contingency tables, 175-177 percentage generation, 155-158
correlation-regression analysis, 479-483 Temporal sequence, 23-24
definition, 175 Test-retest reliability, 76-77
measures Of association, 380 Tests of statistical significance
multiple regression, 540, 543, 544 See Significance tests
partial correlations, 538-541 Theory, 16-18
stepwise multiple regression, 547, 549-552 Theory development, 16-18
Index >» 609

Total sum of squares (SS,,), 327-328 contingency tables, 150-158


Trimodal distribution, 103 continuous variables, 102-103
Tschuprow’s 7; 430 control variables, 164-174, 511-515
t tests criterion variables, 25
alternative formulas, 252-253 definition, 9-10, 455-456
compared to F (F ratio), 337-338 demographic variables, 47-48, 64-65
computer programs, 276, 288-303 dependent variables, 24-25, 172-174
critical values, 248-252, 286 functions, 451-452
definition, 247 4 Goodman and Kruskal’s lambda (A),
dependent samples, 274-275, 297-303, 310 371-377
equality of means, 288-290 grouped data, 151-155
F test for homogeneity of variance, independent variables, 24-25, 172-174
276-297, 309 index and scale construction, 69-73
independent samples, 275-286, interaction effects, 351
290=297,,309 intervening variables, 173-174, 510-511
levels of significance, 282 linear equations, 455-456
one-sample ¢ tests, 201, 246-249 multiple regression, 520-561
paired difference ¢ tests, 297-301 n-by-n tables, 367-377, 386
procedures, 276-277 operational definitions, 65-69
significance tests, 201, 202, 263 predictor variables, 25
standard deviation (0), 246-249, 275-276 relationships, 18-22, 150-174
two-sample ¢ tests, 202, 271-314 two-by-two tables, 360-306, 385-386
Tukey’s honestly significant difference two-category variables, 39-40
test, 343 See also Analysis of variance (ANOVA);
Two-by-two tables, 360-366, 385-386 Correlation-regression analysis;
Two-category variables, 39-40 Measurement; Significance tests
Two-sample ¢ tests, 202, 271-314 Variance, 132-142, 199, 216, 276-297, 310
Two-tailed tests See also Analysis of variance (ANOVA)
See Directional alternative hypothesis
Two-way analysis of variance (two-way ANOVA), Within-groups degrees of freedom (df,,),
348-352 329-330
Type I error, 207—208, 307 Within-groups mean square (MS\,), 329-330
Type II error, 223, 307-308 Within-groups sum of squares (SS), 328
Working definitions, 65-69, 77-79
U
nequal population variance ¢ test, 276-277,
283-287, 289, 310 x-axis, 100-101, 447-451
U ngrouped frequency distribution, 49-50, 86, x-bar (X), 86-90
101-102 x-coordinate, 449-450
nimodal distribution, 103, 107, 111
nit of analysis, 25-26 Yates’s correction for continuity, 415-417
nreliability
aceite y-axis, 447-451
See Reliability y-bar (7), 86
y-coordinate, 449-450
Validity y-hat (predicted value of y), 467-468
chi-square test, 417-421 y-intercept, 456-457
construct validity, 74 Yule’s Q, 362-305
content validity, 73
criterion validity, 73-74 z-bar (Z), 86
definition, 73 Zero, 46-47
face validity, 73 Zero-order correlations, 512-515, 542-543
Variability Zero-order regression slope, 521-522
See Dispersion z tests
Variables origins, 226-229
absolute zero, 46-47 proportions, 253-254, 263
antecedent variables, 173-174 sampling distribution of sample means,
causal models, 22-25 235-246
About the Author

R. Mark Sirkin, Ph.D., is Associate Professor of Political Science at Wright


State University in Dayton, Ohio. He is also the director of the Liberal
Studies degree program in Wright State’s College of Liberal Arts. He teaches
courses On quantitative methods, empirical analysis, and statistical applica-
tions in the social sciences, as well as courses on the Middle East. In the
past, in addition to teaching, he served as Assistant Dean in the College of
Liberal Arts at Wright State University. Among other responsibilities, he
undertook studies and provided data analysis on such subjects as course
enrollment patterns, grade distributions, and other topics pertaining to the
college’s students and faculty.
After that, he was Assistant Dean and later Associate Dean of Wright
State’s School of Graduate Studies. As Associate Dean, he was also Director
of the Office of University Research Services, which administered spon-
sored programs—grants and contracts—as well as supervising the use of
human subjects in research. During that period, he prepared an accredita-
tion self-study for the university and was involved in the analysis of data on
graduate students and faculty research throughout the university, includ-
ing its School of Medicine and School of Professional Psychology. He also
administered the Institutional Review Board for human subjects in
research. These positions gave him a great deal of experience in academic
research, not only in the social and behavioral sciences, but also in the
biological, physical, and medical sciences.
Professor Sirkin earned his B.A. from the University of Maryland and his
M.A. and Ph.D. in political science at Pennsylvania State University.

610 «
An Instructor's CD-ROM containing
data sets, PowerPoint® slides,
exercises, and answers will be
available free of charge to
professors adopting this text.

~ SCIENCES... Aas
Do your aaa lack cae |in ae ability to handle quantitative
work? Do they get confused about how to enter statistical data in SAS®,
SPSS®, and Excel® programs? The new Third Edition of the best-selling
edStatisticsfortheSocial Sciences iis thesolution to these dilemmas.

Lie; previous editions, be Third Edition. continues to help build


students’ confidence and ability in
i doing statistical analysis by slowly
moving {from concepts that require littlecomputational work to ne that
require more. Author [Link] Sirkin «once again demonstrates how
statistics can be used to help students come to appreciate real-world
applications rather than fearing them.‘Statistics for the Social Sciences
emphasizes the analysis and iinterpretation of data to give students a feel
or how data interpretation iis related tothe methods by which the
information was obtained. The book includes lists of key concepts,
— exercises, e ae and more.

New to the Third Edition:


* Includes additional exercises to reflectjhe new computer coverage
° Provides new [Link] teach students how to do analysis not only
through SAS® and SPSS* but also » using Excel® descriptive
statistics features =
* Offers a greater range of othe bea various fields in
i the social
sciences to demonstrate the role ofstatistical analysis in the
=e processee : =<

Statistics for the Social Sciences i


is ansetclieni text for advanced
re undergraduate and graduate students ‘studying statistics across the social
sciences. It can also be used in research methods courses that cover
—: eee in some depth. ;

ISBN 1-4125-O054b-X

II ti iit
91781412 le he I|\||

You might also like