A new questionnaire has just been written. Order the checks a psychometrician runs on it, from the one that must pass first to the one that can only be answered last.
- Do the items agree with each other, or is the scale measuring several unrelated things?
- Do the findings hold in a different population and setting from the one it was built on?
- Does the scale correlate with established measures of the same construct?
- Do the same people score similarly when retested a fortnight later?
- Does it distinguish people already diagnosed from people who are not?
Hints
- The first two checks compare the measure to itself; the rest compare it to the world.
- A scale that gives a different answer every time cannot be accurate, so that has to be settled first.
Show the answer
- Do the items agree with each other, or is the scale measuring several unrelated things?
- Do the same people score similarly when retested a fortnight later?
- Does the scale correlate with established measures of the same construct?
- Does it distinguish people already diagnosed from people who are not?
- Do the findings hold in a different population and setting from the one it was built on?
The order is not arbitrary - it follows the dependency between the concepts. Internal consistency and test-retest stability come first because they are prerequisites: a scale whose items disagree with each other, or whose scores drift between fortnights, cannot be tracking anything stable, and there is no point asking what it measures until it measures something.
Once it is reliable, the validity questions begin, and they compare the scale to something outside itself. Agreement with established measures and the ability to separate diagnosed from undiagnosed groups are both evidence about the construct - the questions the reliability statistics could not answer.
External validity comes last because it needs everything before it plus a second population. A scale validated entirely on undergraduates may behave quite differently in a clinical sample or another culture, and that is a question no amount of work on the original sample can settle.
Practise Reliability and Validity
The app has 6 more questions on this lesson, and keeps your place in the course. Psychology I is free to start.
More questions on Reliability and Validity
- A scale reading 73.0, 73.1, 72.9 and 73.0 for a certified 70 kg weight is reliable but not valid.
- A bathroom scale reading 3 kg heavy every time is perfectly reliable and not valid at all. Why can…
- A depression questionnaire has excellent internal consistency and excellent test-retest reliability. A critic…
- Select every statement that is TRUE.
- Before you can study happiness you must decide what counts as happiness, a rating scale, hours smiling, a…