Measuring Consistently, and Measuring the Right Thing
Two distinct standards determine whether a psychological measure is trustworthy. Reliability is about CONSISTENCY: does the measure give the same result under the same conditions? Test-retest reliability checks whether the same people score similarly when measured again later; inter-rater reliability checks whether independent raters agree; internal consistency (often measured with a statistic called Cronbach's alpha) checks whether different items on the same scale correlate with each other, as they should if they are all measuring the same underlying construct.
Validity is about ACCURACY: does the measure actually capture what it claims to capture? Construct validity asks whether a measure truly captures the underlying theoretical concept (does this "intelligence test" really measure intelligence, or something else, like test-taking speed?). Internal validity asks whether a STUDY's design justifies its causal claims (are confounds ruled out?). External validity asks whether a study's findings generalize beyond the specific sample and setting studied.
| Concept | Question it answers | Example check |
|---|---|---|
| Reliability | Is it consistent? | Test-retest, inter-rater agreement |
| Construct validity | Does it measure the right concept? | Does an "anxiety" scale really tap anxiety? |
| Internal validity | Does the STUDY justify its causal claim? | Are confounds ruled out? |
| External validity | Do findings generalize? | Beyond this sample and setting? |
Reliability and validity are related but distinct: a measure can be highly RELIABLE without being VALID — a bathroom scale that consistently reads 5 pounds too heavy is perfectly reliable (consistent) but not valid (accurate). Validity, however, REQUIRES reliability: an inconsistent measure cannot possibly be capturing the true underlying concept accurately every time, since its readings keep changing for no real reason.
Common pitfall: treating "reliable" and "valid" as synonyms, or assuming a reliable measure must therefore also be valid. Consistency (reliability) is necessary but NOT sufficient for accuracy (validity) — a measure can be perfectly, unfailingly consistent while consistently measuring the wrong thing.
A dartboard: a tight cluster of darts far from the bullseye, labelled 'reliable, not valid'; darts scattered widely but centered around the bullseye on average, labelled 'valid on average, not reliable'.