PSYC 302 Disposition – a characteristic which is fairly stable over time and fairly stable
from situation to situation.
Lecture 2 Personality traits: The manner in which people behave; their personal
Conceptual Variables: Any concept that we can describe but can’t style.
directly measure. Ability traits: level of performance on unfamiliar tasks or those which
have not involve training.
We need to represent CV with MV through operationalization.
STATES: Short-lived; depend greatly on life events, thoughts, or physiological
Operational Definition: Researchers specific decision about how to effects.
measure a conceptual variable.
ATTITUDES: A set of emotions, beliefs, and behaviors toward a particular
MEASUREMENT SCALES object, person, thing, or event
Scaling: Specifying relationship btw the values of MV-CV.
Domain: The range of possible items which could be measured.
Nominal: Named (gender/ yes*no)
Ordinal: Named, ordered (likert scale) In history class, there is 900 historical dates. In this example the
domain of item is 900 historical dates.
Interval: Named, ordered, equal interval (thermometer, GRE exam
True Score: The aim of any scale is to estimate a person’s true score.
score)
There is no estimation in this approach: Their score on all of the items in the
Ratio: Named, ordered, equal interval, has “zero” (length, weight) } domain is their true score.
Ratio btw values can be calculated
Items need to be representative with respect to relevant dimensions.
(validity issue)
ITEM, SCALES, TESTS
Stratified random sampling involves the division of a population into smaller
What do psychological tests measures? groups known as strata.
Traits (Personality, ability) Strata is formed based on shared attributes or characteristics.
States (Mood, motivation)
Attitudes Random samples selected from each strata.
Single Item Scales (when domain is narrow) / Multiple Item Scales
TRAIT
Types of items measuring personality, mood, motivation and attitude Lecture 3
Projective tests (personality or psychopathology) Psychological measurement measures the attributes of persons,
‘Objective tests’ (unable to modify behavior)
Researcher`s roadmap in the absence of a well-formulated attribute
Rating scales – We will focus primarily on this in this course (used to
assess attitude)
theory
Response Bias Step 1: Identify an attribute of interest (genuine characteristic*smt
inside*, not socially constructed) (Proneness to guilt)
Acquiescent responding: Tendency to select a positive response - Genuine attribute vs. imputation (beauty)
- Genuine attribute vs social construction
Scale with reverse coded/keyed items to understand A.R. bias.
GA should be reflective (not formative) which is causal factor
Self-Report Rating Scales explaining the behavior. Indicators correlative.
Very popular because: Step 2: Decide on your research approach
- easy to apply and understand - Option 2: Data-based measurement (Skip to Step 4)
- very flexible in the type of questions -
Potential problems: “Situations can be ordered in terms of their “guilt-inducing power”
- Assumes that respondents understand the questions & (δ)
understand them in the same manner “People can be ordered in terms of their “guilt-sensitivity” (θ).
- Relies on respondents` own evaluations of their feelings, So, if a situation i dominates person j, δi > θj , person j will feel guilty
thoughts, behaviors. in situation i.
- Assumes that people are willing to report honestly. The intensity of guilt depends on the magnitude δi – θj.
- Reactivity (Participants may alter their behaviour because they Operationalize the attribute into a set of items.
know they are part of a study) Items need to represent varying levels of guilt-inducing power.
Reactivity as a limitation to self-report
measures Step 3: Conduct studies to understand the underlying structure of
Social desirability bias– tendency to present the attribute, or underlying mechanisms that give rise to
yourself in the socially acceptable way. observables (item scores)
Confusing question Step 4: Put together an experimental test version of the scale and
Leading Question collect data
Negatively worded questions - Items with varying degrees of guilt-inducing power.
Double-barreled questions (2+ questions, but allows only 1 - Item response function (IRF) consistent with Guttman
response.) (also can be in response) model.
Guttman scale is the technique that assess the extent of the subject’s
agreement with items, where the items are arranged in ordered.
It is deterministic, rather than probabilistic.
Assumes no error in data. In more realistic data there will be errors, that’s
why it is hard to form Guttman scales in psychological testing.
Step 5: Psychometric analysis (Conduct model-based analysis (factor
Latent Ability(guilt proneness)
analysis, IRT))
Step 6: Feedback to the theory of the attribute
Statistical measurement models
IRT is measurement model Lecture 4
IRT model estimates the likelihood of different responses to items with different
levels of trait being measured.
“Theta” is a symbol that stands for the trait being measured. Probability of
IRT define the estimated probability of a given response to given item . responding yes to item i.
Item response theory (IRT), also called latent trait theory, is a psychometric Theta is trait level (level of pain)
theory that was created to better understand how individuals respond to
items on psychological tests. How likely a person to pick a particular item response depends on
The term latent trait is used to describe IRT in that characteristics of how much of the trait they
individuals which cannot be directly observed. have how difficult the item is.
The x-axis relates to the level of the characteristic being measured. Mapping between item/scale scores and true attribute level depends on the
extent to which we can control the factors that determine a particular item
This trait, is typically scored like a z score, with scores at zero being average response.
and scores above zero being above average; scores below zero are below
average. Perfect control of the factors** deterministic (exact) mapping
Less than perfect control** probabilistic mapping
Success of this mapping depends on some properties of the items that make
up the scale. ( Item difficulty, Item discrimination)
For example, we might expect an exam composed of difficult items to
do a great job in differentiating top respondents, but it is worthless
for the lower half of respondents because they will be so confused
and lost.
Dichotomously-scored item
It includes two possible item scores.
Multiple choice, 4 to 5 options, but only two possible scores
(correct/incorrect).
Trace line(item characteristic curve, item characteristic function)
Polytomously-scored item
(yellow) = probability of response ‘yes
Polytomous models are for items that have more than two possible
scores.
Likert-type items (Rate on a scale of 1 to 5) and partial credit items
(score on an Essay might be 0 to 5 points).
ITEM PROPERTIES
Item difficulty refers to how easy or difficult each item is for our sample.
The Item information function is a measure of how much statistical
information a test item provides.
Item difficulty and item discrimination are the two item characteristic
Item discrimination refers to the ability of an item to differentiate among that are estimated in Item Response Theory (these are called item
students based on how well they know the material being tested. parameters)
The Test Information Function how well an assessment differentiates
respondents, and at what ranges of ability.
Item discrimination represented by slope (eğim)
Polytomously Item
We have more than two probability two items represented by trace line are different in diffuculty but they
are equal in discrimination.
Two items have same difficulty. (median probabilityden çizgi çek) but
they have different discrimination.
Purple has higher slope, more discriminative. (it discriminate among
people who are similar levels of trait being measured)
More changes in trait level equal very big changes in probability of
endorsing item. It means that it helps you discriminate btw person
who are close to same level of trait but not same level of trait.
Item with higher discrimination have higher information.
The greatest information is where the probabilities changing a lot the
most likely response category changing a lot.
Item represent by the purple one is easier, easier to say yes.
Yellow item is difficult than the purple. All of These are used for decide on which item to include.
Discrimination = 0 } item is irrelevant to the attribute being
measured. Line will be ----------------
“Item discrimination tells how closely related an item is to
attribute.
Always screen for negative discrimination.
o Could be due to a reverse-scored item that went
unnoticed.
o Could be due to a poorly written item.
Lecture 5
Once we estimate item properties, how do we use this information when
developing the scale?
Item calibration is a technique to estimate characteristics of questions (called
items) for achievement tests. **How do we determine/estimate item
parameters?
In computerized tests, item calibration is an important tool for maintaining,
updating and developing new items for an item bank.(?)
Test calibration is accomplished by administering a test to a group of M
examinees and dichotomously scoring the examinees’responses to the N
items. Then mathematical procedures are applied to the item response data
in order to create an ability scale that is unique to the particular combination
of test items and examinees. Then the values of the item parameter
estimates and the examinees’ estimated abilities are expressed in this metric.
Once this is accomplished, the test has been calibrated, and the test results
can be interpreted via the constructs of item response theory
Lecture 6
Category thresholds(?))
Lecture 7
Classical definition of measurement