0% found this document useful (0 votes)
6 views20 pages

Understanding Psychological Testing Types

The document outlines the purpose and classification of psychological testing, including its applications in identifying problems such as brain damage, learning disabilities, and emotional disorders. It discusses various types of tests, including cognitive, non-cognitive, and personality tests, as well as issues of bias, fairness, and ethical considerations in psychological testing. Additionally, it emphasizes the importance of norms and scoring interpretation for evaluating test results and ensuring proper use of psychological assessments.

Uploaded by

gargvanshika715
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views20 pages

Understanding Psychological Testing Types

The document outlines the purpose and classification of psychological testing, including its applications in identifying problems such as brain damage, learning disabilities, and emotional disorders. It discusses various types of tests, including cognitive, non-cognitive, and personality tests, as well as issues of bias, fairness, and ethical considerations in psychological testing. Additionally, it emphasizes the importance of norms and scoring interpretation for evaluating test results and ensuring proper use of psychological assessments.

Uploaded by

gargvanshika715
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Module I Introduction

Purpose of testing,
Classification of people-
assigning a person to a particular category for different purposes  (selection
and classification of industrial personnel, military personnel)
placement, - sorting into programs based on appropriate skills
screening, - quick and simple tests to identify people who might have special or
required characteristics or skills
certification – have a pass/fail quality  gives privilege to attend uni or get a job

Identifying problems
Brain damage:
Neuropsych tests are used in the assessment of individuals with known or
suspected brain damage.
Learning disability:
identification of slow, fast learners to identify children with learning disabilities
Disorders (emotional/behavioural):
clinical tests include examination of severe emotional disorders and other types
of behavioural problems
Intellectual Deficiencies: detection of intellectual deficiencies eg. Binet tests

Evaluation
Self knowledge- feedback
Program evaluation-
evaluate impact and effectiveness of social programs and policies.
Same person, different conditions:
tests are used to measure differences between individuals or reactions of the
same individual under different conditions
Counselling and Therapy
Career or vocational counselling
Diagnosis treatment/intervention planning-

Research-
Evaluate existing theories
test and explore new hypotheses.

types of test used,


Types of tests used

It is important to note that there is no correct cataloguing of the types of tests as


the different categorizations often overlap.
Clinical Structured interview
Observation (behavioural procedures)
objectively describe and count frequency of behaviour  identify antecedents
and consequences
Direct
Indirect

Record review, MSE, Case history, custom-questionnaires etc


Psychological tests
Neuropsychological tests
measure cognitive, sensory, perceptual and motor performance to determine
extent, locus and behavioural consequences of brain damage
Cognitive tests (performance validity)
1. maximal behaviour (intelligence and ability measures)
2. Performance validity tests check if person is putting in sufficient
effort to perform to their best capability (cognitive test).

Intelligence
measures an individual’s ability in global areas such as verbal comprehension,
perceptual organisation, or reasoning- help determine potential for scholastic
and other performance
Intelligence tests, like aptitude tests, assess the test taker’s ability to cope with
the environment, but at a broader level. Intelligence tests are often used to
screen individuals for specific programs (for example, gifted programs, honors
programs) or programs for the mentally challenged. Intelligence tests are
typically used in educational and clinical settings.

Aptitude
measures capability for certain tasks or skills (how much they CAN grow)
Aptitude tests assess a test taker’s potential for learning or ability to perform in a
new job or situation. Aptitude tests measure the product of cumulative life
experiences—or what one has acquired over time. They help determine what
“maximum” can be expected from a person.
Achievement
measures a persons degree of learning, success (how much they HAVE grown)
Achievement tests measure a person’s previous learning in a specific academic
area (for example, computer programming, German, trigonometry, psychology).
A test that requires you to list the three characteristics of psychological tests
would be considered an achievement test. Achievement tests are also referred to
as tests of knowledge.
Memory
Language
Attention
Creativity-
assess novel or original thinking and capacity to come up with unusual solutions
for problems

Non-cognitive tests (symptom validity)


1. typical behaviour (personality, interests, values, and attitudes).
2. Commonly self report.
3. Structured (yes no Q like do you engage in this activity or no) or
4. Unstructured (personalty tests- projective techniques; Rorschach; scoring
is much more complex). Projective tests- looks at underlying and
unconscious motivations and attitudes through ambiguous stimuli.
5. Symptom validity testsproviding accurate self report (non cog)
Personality
measures traits, qualities or behaviours that determine a person’s individuality
Personality tests measure human character or disposition. The first personality
tests were designed to assess and predict clinical [Link] tests remain
useful today for determining who needs counseling and who will benefit from
treatment programs. Newer personality tests measure “normal” personality
traits. For example, the Myers–Briggs Type Indicator (MBTI) is often used by
industrial/organizational psychologists to increase employees’ understanding of
individual differences and to promote better communication between members
of work teams. Career counselors also use the MBTI to help students select
majors and careers consistent with their personalities.

Objective
Projective
Interest
measures an individual’s preference for certain activities helps in determining
occupational choice
Interest inventories assess a person’s interests in educational programs for job
settings and provide information for making career decisions. Because these
tests are often used to predict satisfaction in a particular academic area or
employment setting,they are administered primarily to students by counselors in
high schools and colleges

Cognitive tests Non Cognitive


Maximal behaviour- 2 types- Measures Personality, vocation
ability (learned from or occupational interests
environment) and achievement symptom validity measures
(learned from institutions) tests.
Both tests involve learning. **not ability tests. Provide more
*measure what people know and of qualitative descriptions of
can do (achievement) or broad people. Most are self report or
areas of what they are capable self descriptions tools
of doing (ability; verbal
reasoning, numerical reasoning,
spaitial reasoning)
Ability tests- verbal + Measures of Drive, Motivations
performance (most commonly and Need.
used for intelligence tests). Usually administered without
Usually timed. time limit.
best predictor of intelligence test Personality inventories: looks at
performance is one's vocabulary (learned the way we respond to other
at home, school and social environment; people and situations. “tests that
mix of ability and achievment), which is assess disposition (preferred way
why it is often given as the first test of thinking). No right or wrong
during intelligence testing or in some ans.
cases represents the body of the
intelligence test (e.g., the Peabody Picture Interest inventories: also related
Vocabulary Test to personality, but focus on
activities that are found
attractive. Obvious utility in staff
development assessment
situations at work. Help explore
new options for people.
Speeded vs power tests.
Speeded- one that everyone
would score if they had enough
time (reaction time, attention).
Power- influenced by how much
the individual knows
.
Can be free response (fill in the a recognition question on a non-
blanks)or recognition (MCQ). cognitive test might ask
Also observe the process someone whether they would
through which the individual rather go ice skating or to a
responses as opposed to just movie; a free recall question
answer. would ask the respondent what
***can be speed tests. they like to do for enjoyment.
Bias & Fairness
Bias
Presence of Systematic error
Bias in testing refers to the presence of systematic error in the measurement of
certain factors (e.g., academic potential, intelligence, psychopathology) among
certain individuals or groups (Sandoval et al., 1999; Suzuki, Meller, & Ponterotto,
1996).
Disproportionate for minorities
The basic premise is that a screening device (psychological test) can have an
adverse impact if it screens out a proportionally larger number of minorities than
nonminorities.

Difference should be real not cultural


This conclusion is based largely on early intuitive observations that many African
American children and other minorities usually do not have the opportunity to
learn the types of material contained in many of the test items. Thus, their lower
scores may represent not a lack of intelligence, but merely a lack of familiarity
with European American, middle-class culture
Tests are intended to discriminate between people – to show up differences
where these are real. What they should not do is discriminate unfairly. That is,
show differences where none exist, or fail 1to show differences that do exist. It is
possible that factors such as sex, ethnicity or social class may act to obscure,
mask or bias a person’s true score on a test. It is bias that we need to remove or
minimise in the design of tests, not differences.

Fairness (moral, philosophical, legal issue)


Thorndike (1971) wrote, “The presence (or absence) of differences in mean score
between groups, or of differences in variability, tells us nothing directly about
fairness” (p. 64). In fact, the concepts of test bias and unfairness are distinct in
themselves.
A test may have very little bias, but a clinician could still use it unfairly to
minority examinees’ disadvantage. Conversely, a test may be biased, but
clinicians need not—and must not—use it to unfairly penalize minorities or others
whose scores may be affected.
Jensen (1980) was the author who first argued cogently that fairness and bias
are separable concepts. As noted by Brown et al. (1999), fairness is a moral,
philosophical, or legal issue on which reasonable people can legitimately
disagree. By contrast, bias is an empirical property of a test, as used with two or
more specified groups. Thus, bias is a statistically estimated quantity rather than
a principle established through debate and opinion
(Graham & Jack, 2003, p. 71)
1. Unfairness defi ned. When members of one race, sex, or ethnic group
characteristically obtain lower scores on a selection procedure than
members of another group and the diff erences in scores are not refl ected
in diff erences in a measure of job performance, use of the selection
procedure may unfairly deny opportunities to members of the group that
obtains the lower scores.
2. Investigation of fairness. The greater the severity of the adverse impact on
a group, the greater the need to investigate the possible existence of
unfairness.
3. When unfairness is shown. If unfairness is demonstrated through a
showing that members of a particular group perform better or poorer on
the job than their scores on the selection procedure would indicate
through comparison with how members of other groups perform, the user
may either revise or replace the selection instrument in accordance with
these guidelines, or may continue to use the selection instrument
operationally with appropriate revisions in its use to ensure compatibility
between the probability of successful job performance and the probability
of being selected.
4. Continued use of selection procedures when fairness studies not feasible.
If a study of fairness should otherwise be performed, but is not technically
feasible, a selection procedure may be used which has otherwise met the
validity standards of these guidelines, unless the technical infeasibility
resulted from discriminatory employment practices which are
demonstrated by facts other than past failure to conform with
requirements for validation of selection procedures. However, when it
becomes technically feasible for the user to perform a study of fairness
and such a study is otherwise called for, the user should conduct the study
of fairness.
(Kaplan & Sacuzzo, 2009, pp. 552-553)

Ethical Issues in Psychological Testing


The ITC has produced international guidelines on test use (available from their
website: [Link]) that have been endorsed by the Society. These
guidelines embody the same principles of good practice that the Society has
embedded within its test user qualifications and its Code of Good Practice in
Psychological Testing (see Appendix C).
These various codes are based on some very simple common-sense principles:
1. You should know the limits of your own competence.
2. You should be competent in what you do.
3. You should know the strengths and limitations of the tools you use.
4. You should treat all people involved in the testing process with respect.
5. You should ensure that you have their informed consent to the test
conditions.

Appendix C: British Psychological Society Code of Good Practice


for Psychological Testing

RESPONSIBILITY FOR COMPETENCE


People who use psychological tests are expected by the British Psychological
Society to:
1. Take steps to ensure that they are able to meet all the standards of
competence defined by the Society for the relevant Certificate(s) of
Competence in Testing, and to endeavour, where possible, to develop and
enhance their competence as test users.
2. Monitor the limits of their competence in psychometric testing and not to
offer services, which lie outside their competence nor encourage or cause
others to do so.
3. Ensure that they have undertaken any mandatory training and that they
have the specific knowledge and skills required for each of the
instruments they use.
PROCEDURES AND TECHNIQUES
People who use psychological tests are expected by the British Psychological
Society to:
4. Use tests, in conjunction with other assessment methods, only when their
use can be supported by the available technical information.
5. Administer, score and interpret tests in accordance with the instructions
provided by the test distributor and to the standards defined by the
Society.
6. Store test materials securely and to ensure that no unqualified person has
access to them.
7. Keep test results securely, in a form suitable for developing norms,
validation, and monitoring for bias.
CLIENT WELFARE
People who use psychological tests are expected by the British Psychological
Society to:
8. Obtain the informed consent of potential test takers, making sure that
they understand why the tests will be used, what will be done with their
results and who will be provided with access to them.
9. Ensure that all test takers are well informed and well prepared for the test
session, and that all have had access to practice or familiarisation
materials where appropriate.
[Link] due consideration to factors such as gender, ethnicity, age, disability
and special needs, educational background and level of ability in using
and interpreting the results of tests.
[Link] the test taker and other authorised persons with feedback about
the results in a form, which makes clear the implications of the results, is
clear and in a style appropriate to their level of understanding.
[Link] test results are stored securely, are not accessible to unauthorised
or unqualified persons and are not used for any purposes other than those
agreed with the test taker.

Overview of Tests

Norms, Scoring Interpretation


1. (comparative frame of reference) norm stands for the normal or
average performance, to which the score of the individuals performance is
compared to make inferences.
2. In the process of standardization, the test is conducted on a large,
representative sample of the category of people for whom the test is
designed. Analysis of the general performance then decides how
performance of person will be predicted. (establish average
performance)
3. Interpretation Scores (e.g. 16 out of 25 items correct) obtained on tests
are typically converted into a ‘standard’ form to facilitate their
interpretation.
4. This may be carried out by using tables of ‘norms’ or by reference to
criterion scores. Norms provide information about the distribution of
scores in some population (for example, ‘UK working adults’) and scores
can be converted into numbers that show how a person has performed
relative to this population.
5. For e.g. Instead of saying the person got 16 out of 25 correct, we might
say they performed at a level equivalent to the top 30 per cent of the UK
working adult population.
6. Norms are important because the latter type of statement is more
meaningful and useful than the former.
7. A much more powerful approach is to use the relationship between test
scores and criterion measures. These are external measures of interest,
such as training outcome, job success, categories of mental dysfunction,
etc. Criterion measures provide another means of aiding the interpretation
of scores. To take a very simple example, if we know (from our validation
research) that the failure rate in a training course of 50 per cent for people
who score less than 10 on a test, 35 per cent for those who score between
10 and 15, and only 20 per cent for those who score 16 or above, then we
can criterion-reference the score by converting the scores into predicted
training outcomes. In effect we can classify the people on the basis of their
test scores in terms of risk of training failure.
(psychological testing: a user’s guide)
*** Norms also indicate frequency with which high and low score are obtained 
show the degree to which a score deviates from expectations or the MEAN
(average).
Percentile scores- that describes where a person matched with variables like age
and/or gender stands with respect to the sample population used to standardise
the test. This is used for interpreting a test by comparing performance of
individual to a comparable experiential background.

Percentage score- is basically the raw score that determines the percentage of
correct answers in the test.

Utility of a test is also determined by the consistency and repeatability of the


test – reliability
- validity ( accurations of interpreting)

initial outcome of test is a raw score, which is meaningless by itself, unless


converted into a derived score based on comparison (composite score,
percentile rank etc) standardisation or norm group (sample of examinees
representative of the population for whom the test is intended).

To analyse norms

 Establish Mean (avg) , median (central most score) and Mode (highest
frequency score) look at frequency of score distribution. Normal
distribution curve – bell shaped (no kurtosis- mesokurtic, asymptotic
curve), Mean= median= mode), Skewness=0
Assuming normal distribution is useful because, a) it helps accurately
calculate area under the curve and b) calculate % of scores that fall
within a range of scores. Eg., 68% of the scores will fall within 1SD of
mean from both directions.
Skewness- symmetry of the curve- frequency of scores in which
direction. Eg., if scores are piled up at the high end of the scale
(negative skew)= scale is too easy or low end (positive skew)= too
hard.
** Skewness is corrected by changing or adding test items to correct
for level of difficulty. If its too late, to change the items, can use
statistical transformation. But preferred method is to modify test items
to minimise skewness.
Standard Deviation- degree of deviation from mean.

Selecting a norm group


obtain a representative cross section of population for whom the test is
intended.

Age norms- depicts level of performance for each separate age group. To
facilitate same age comparisons. Interval grouping of age will depend on what is
being measured and how quickly these abilities change through the course of
development.

Grade norms – for each separate academic year. Useful in school settings.

Local norms- representative local sample as opposed to national sample.


Sub group norms – based on ethnic groups or females or just generally groups
with same characteristics (sex, ethnic group, geographical region, urban vs rural,
socioeconomic etc)

Raw score transformation


 transforming raw scores into standard scores doesn’t change the shape of the
distribution.

Percentiles and Percentile Ranks

Percentile- percentage of individuals below a specific raw score in the


standardised sample. Most common raw score transformation.
Percentage- denotes the number of correct responses in a test
Percentile rank- *opp of normal ranking procedures, here, 1 is for the lowest and
100 for the highest. Rank of 50 is for the median or middlemost score. Percentile
of 25 (denoted as Q1 or first quartile) as 25% or 1 quarter of the scores fall below
this point.

*******but percentiles can be misleading, as the difference in the raw scores


obtained in the same interval of percentile but at diff levels (for eg, 50-59 and
90-99) will be very varied.
******Specially in Skewed or NOT NORMAL distribution
Standard Scores (Z score) – (most desirable psychometric properties)

Standard score expresses the distance from the mean in standard


deviation units. For eg. A raw score exactly 1SD above the mean will get
the standard score of 1.
***the standard score always retains the relative magnitude and
differences found in raw scores.
Calculate it by subtracting raw score from M and dividing by the SD.
***good for comparing performance on different tests by converting it into
performance as compared to the average performance on the specific
tests. ******only when the distribution is the same. (otherwise, same
standard score will depict different percentiles, if there is high kurosis-
leptokurtic graph and mesokurtic graph, frequency of scored within 1sd
and getting Z score of 1 will be higher than at mesokurtic. **** non linear
transformation of the raw scores
Greater the Z score, the farther it is from the mean (along x Axis; so at
higher percentile)
*Use the percentile for each raw score to convert into standard score
***too many decimal points and negative sign can be distracting or
confusing***

-3SD  2.14% cases


-2SD13.59% cases
1SD 33.14% cases

T scores and other standardised scores


Show the same information as standard scores
Standardised scores always expressed as whole numbers. (eliminate
fractions and negative sign by producing value other than 0 for the mean
and 1 for sd of the transformed scores). They tailor any M and SD to
eliminate negative and fractions.
T test is a common standardised score where M=50, SD=10 of the
converted scores.
*very common in personality tests.
T score= 10 (X-M) +50

SD
Or T= 10(Z score) + 50

Stanine, Sten and C Scores

Stanine (standard 9 ) scale, developed by US air force, where all raw


scores are converted into scores ranging from 1-9.
M= 5, SD=2
*scores are ranked from lowest to highest, then bottom 4% get score of 1,
next 7% get score of 2, next 12% get a score of 3….

10unit scale Sten scale


C scale 11 units. Stanines very commonly used.

Report Writing
Subject information
Name:
Age (date of birth):
Sex:
Ethnicity:
Date of Report:
Name of Examiner:
Referred by:

Referral Question
The Referral Question section provides a brief description of the client and a
statement of the general reason for conducting the evaluation.
In particular, this should include a brief description of the nature of the problem.
If this section is adequately completed, it should give an initial focus to the
report by orienting the reader to what follows and to the types of issues that are
addressed.

Evaluation Procedures
The report section that deals with evaluation procedures simply lists the tests
and other evaluation procedures used but does not include the actual test
results. Usually, full test names are included along with their abbreviations. Later
in the report, the abbreviations can be used, but the initial inclusion of the entire
name provides a reference for readers who may not be familiar with test
abbreviation.

Behavioral Observations
A description of the client’s behaviors can provide insight into his or her problem
and may be a significant source of data to confirm, modify, or question the test-
related interpretations. These observations can be related to a client’s
appearance, general behavioral observations, or examiner-client interaction.
Descriptions should be tied to specific behaviors and should not represent a
clinician’s inferences. For example, instead of making the inference that the
client was “depressed,” it is preferable to state that “her speech was slow and
she frequently made self-critical statements such as ‘I knew I couldn’t get that
one right.’”

Background Information (relevant history)


The write-up of a client’s background information should include aspects of the
person’s history that are relevant to the problem the person is confronting and to
the interpretation of the test results.
When describing a client’s background, it is important to specify where the
information came from (“The client reported that . . .”). This is particularly
essential when there may be some question regarding the truth of the client’s
self-reports or when the history has been obtained from multiple sources.
Usually, a history begins with a brief summary of the client’s general
background. This can be followed by sections describing family background,
personal history, medical history, history of the problem, and current life
situation.

Test Results
If actual test scores are included, standard (rather than raw) scores should be
the mode of presentation. Referral sources have consistently indicated that
percentiles are preferred over other types of standard scores (Finn et al., 2001).
Because various tests use somewhat different types of standard scores, it is
recommended that each set of test scores include both the standard score and
percentiles. Clinicians may also wish to indicate the relative magnitude of the
relevant scores (“Very high,” “High,” etc.) or whether the scores exceed some
clinically meaningful cutoff.

Impressions and Interpretations (Discussion)


This section can be considered the main body of the report. It requires that the
main findings of the evaluation be presented in the form of integrated
hypotheses.
All inferences made in the Impressions and Interpretations section should be
based on an integration of the test data, behavioral observations, relevant
history, and additional available data. The conclusions and discussion may relate
to areas such as the client’s overt behavior, self-concept, family background,
intellectual abilities, emotional difficulties, medical disorders, school problems, or
interpersonal conflicts.

Summary and Recommendations


The purpose of the summary subsection is to restate succinctly the primary
findings and conclusions. This requires that the practitioner select only the most
important issues and that he or she be careful not to overwhelm the reader with
needless details. As emphasized previously, a useful strategy in the summary
section is to provide brief bulleted/numbered answers to each of the referral
questions.
The ultimate practical purpose of the report is contained in the recommendations
because they suggest what steps can be taken to solve problems. Such
recommendations should be clear, practical, and obtainable, and should relate
directly to the purpose of the report. The best reports are those that help the
referral sources and/or the clients solve the problems they are facing

Issues in measurement
Measurement Concepts
1. Operational Definition: is the definition of a variable in terms of the
actual procedures used by the researcher to measure and/or manipulate
it.
a. Similar to a ‘recipe,’ operational definitions specify exactly how to
measure and/or manipulate the variables in a study.
b. Good operational definitions define pro-cedures precisely so that
other researchers can replicate the study.
c. Impulsivity was operationalized as the total number of incorrect
stimulus responses
d. Two doses of alcohol were used: 5g/kg and 10g/kg
e. Alcohol dependence vulnerability was defined as the total score on
the Michigan Alcohol Screening Test (MAST; Selzer, 1971)
2. Measurement error
a. A participant’s score on a particular measure consists of 2
components:
b. Observed score = True score + Measurement Error
i. True Score = score that the participant would have obtained
if measurement was perfect—i.e., we were able to measure
without error
ii. Measurement Error = the component of the observed
score that is the result of factors that distort the score from
its true value
3. Factors that Influence Measurement Error
a. Transient states of the participants:
b. (transient mood, health, fatigue-level, etc.)
c. Stable attributes of the participants:
d. (individual differences in intelligence, personality, motivation,
etc.)
e. Situational factors of the research setting:
f. (room temperature, lighting, crowding, etc.)
4. Characteristics of Measures and Manipulations
a. Precision and clarity of operational definitions
b. Training of observers
c. Number of independent observations on which a score is based
(more is better?)
d. Measures that induce fatigue or fear
5. Actual Mistakes
a. Equipment malfunction
b. Errors in recording behaviors by observers
c. Confusing response formats for self-reports
d. Data entry errors
e. (Measurement error undermines the reliability (repeatability) of the
measures we use)

Reliability
• The reliability of a measure is an inverse function of measurement error:
• The more error, the less reliable the measure
• Reliable measures provide consistent measurement from occasion to
occasion

Different methods for assessing reliability


• Test-Retest Reliability
• Test-retest reliability refers to the consistency of participant’s
responses over time (usually a few weeks, why?)
• Assumes the characteristic being measured is stable over time—not
expected to change between test and retest
• Inter-rater Reliability
• If a measurement involves behavioral ratings by an observer/rater,
we would expect consistency among raters for a reliable measure
• Best to use at least 2 independent raters, ‘blind’ to the ratings of
other observers
• Precise operational definitions and well-trained observers improve
inter-rater reliability
• Internal Consistency Reliability
• Relevant for measures that consist of more than 1 item (e.g., total
scores on scales, or when several behavioral observations are used
to obtain a single score)
• Internal consistency refers to inter-item reliability, and assesses the
degree of consistency among the items in a scale, or the different
observations used to derive a score
• Want to be sure that all the items (or observations) are measuring
the same construct
• Estimates of Internal Consistency
• Item-total score consistency
• Split-half reliability: randomly divide items into 2 subsets
and examine the consistency in total scores across the 2
subsets (any drawbacks?)
• Cronbach’s Alpha: conceptually, it is the average
consistency across all possible split-half reliabilities.
Cronbach’s Alpha can be directly computed from data

Validity
1. A good measure must not only be reliable, but also valid
2. A valid measure measures what it is intended to measure
3. Validity is not a property of a measure, but an indication of the extent to
which an assessment measures a particular construct in a particular
context—thus a measure may be valid for one purpose but not another
4. A measure cannot be valid unless it is reliable, but a reliable measure may
not be valid
5. Like reliability, validity is not absolute
6. Validity is the degree to which variability (individual differences) in
participant’s scores on a particular measure, reflect individual differences
in the characteristic or construct we want to measure.

Three types of measuring validity


1. Face Validity
a. Face validity refers to the extent to which a measure ‘appears’ to
measure what it is supposed to measure
b. Not statistical—involves the judgment of the researcher (and the
participants)
c. A measure has face validity—’if people think it does’
d. Just because a measure has face validity does not ensure that it is a
valid measure (and measures lacking face validity can be valid)
2. Construct Validity
a. Most scientific investigations involve hypothetical constructs—
entities that cannot be directly observed but are inferred from
empirical evidence (e.g., intelligence)
b. Construct validity is assessed by studying the relationships between
the measure of a construct and scores on measures of other
constructs
c. We assess construct validity by seeing whether a particular
measure relates as it should to other measures
d. Self-esteem example
i. Scores on a measure of self-esteem should be positively
related to measures of confidence and optimism
ii. But, negatively related to measures of insecurity and anxiety
e. Convergent and Discriminant Validity
f. To have construct validity, a measure should both:
i. Correlate with other measures that it should be related to
(convergent validity)
ii. And, not correlate with measures that it should not correlate
with (discriminant validity)
3. Criterion-related Validity
a. Refers to the extent to which a measure distinguishes participants
on the basis of a particular behavioral criterion
b. The Scholastic Aptitude Test (SAT) is valid to the extent that it
distinguishes between students that do well in college versus those
that do not
c. A valid measure of marital conflict should correlate with behavioral
observations (e.g., number of fights)
d. A valid measure of depressive symptoms should distinguish
between subjects in treatment for depression and those who are not
in treatment
e. Two types of criterion related validity
i. Concurrent validity measure and criterion are assessed at
the same time
ii. Predictive validity elapsed time between the
administration of the measure to be validated and the
criterion is a relatively long period (e.g., months or years).
Predictive validity refers to a measure’s ability to distinguish
participants on a relevant behavioral criterion at some point
in the future
iii. SAT EXAMPLE
1. High school seniors who score high on the SAT are
better prepared for college than low scorers
(concurrent validity)
2. Probably of greater interest to college admissions
administrators, SAT scores predict academic
performance four years later (predictive validity)

Emerging trends of online testing


Ways computer are used in Clinical Assessment
1. Scoring and data analysis
2. Profiling and charting of test results
3. Listing of possible interpretations
4. Evolution of more complex test interpretation and report generation
5. Adapting the administration of test items
6. Decision making by computer

Computer Assisted-Interview
1. In place of the traditional interview/assessment in paper and pencil form
2. Computer asks more standardized questions, covering areas that a
clinician may neglect
3. Facts are recorded more systematically
4. Reduces the likelihood of skewed responses because of social desirability
5. More likely to share sensitive and personal information

Computer-Administered Tests
1. Equivalent to paper-pencil forms for most tests, like MMPI-2, SCII, MAP etc
2. Less-time consuming
3. More cost-effective
4. Drawbacks
a. Computer anxiety may increase negative affect in some particular
tests.
b. Body language won’t be captured by the computer

Computer Diagnosis, Scoring and Reporting of Results


1. Scoring and report generation is straightforward, efficient and accurate.
2. Software has been develop to screen for psychopathology on the basis of
DSM
3. Evaluation of online version of Rorshach showed that computer can
provide a report similar to that of clinician.
4. Errors have been reported too. So far, they can only be used as an
appendage, and not replacement of clinical judgment and expertise.

Internet Usage for psychological Testing


1. A Source of gather data from a large number of participants.
2. Increased the reach of tests in mainstream population, as so many
psychological tests are available for free online. MBTI,
[Link], [Link] (IQ tests)
3. more accessible to people. Reduce barriers like transportation issues.
Reach remote areas where physical forms of such services are not
available.
4. Tests can be administered in more interesting manners.
5. test will be administered in a more consistent and uniform manner, and
avoid variation caused by administration of test by different individuals.
6. offer more widespread awareness, and sensitization to testing, familiarize
people with the process, making the process much more comfortable.
7. Better, faster and cheaper services and products.
a. Test can be downloaded instantly.
b. Revise and update tests within minutes at virtually no cost of
distributing new test forms,
c. Saves time, rapid communication of findings/ results to clients,
researchers and public
8. new test wih translations available all over the world, instantly.
9. Problem:
a. Issue of test or data security
b. provide labels/ diagnosis that should not be given to individuals
without the presence of a trained professional. Lead to
misinterpretation and half information.
c. difficulty in standardizing testing conditions, Internet speed,
Distractions Kinds of computer.
d. Maybe unsuitable for people with disabling conditions
e. Maybe unfit for culturally diverse persons
f. Cheating by respondent, specifically in power tests.
[Link] it’s been shown, results are equivalent to their paper-pencil forms.
Less error in fact.

Computers in Cognitive-Behavioural assessment


Farrel (1992) has identifi ed seven applications of computers in the fi eld of
cognitive-behavioral assessment:
1. (1) collecting self-report data,
2. (2) coding observational data,
3. (3) directly recording behavior,
4. (4) training,
5. (5) organizing and synthesizing behavioral assessment data,
6. (6) analyzing behavioral assessment data, and
7. (7) supporting decision making.
(Kaplan & Sacuzzo, 2009, p. 431)

You might also like