Reliability
Reliability is consistency. A test is reliable when repeated measurements give similar results - across administrations, across raters, across alternate forms, or across items. A reliable test can still be wrong (reliably measuring the wrong thing is possible), but an unreliable test is unusable regardless.
Test-Retest Reliability
Test-retest reliability. The same test is administered to the same subjects on two occasions, and the scores are correlated. High positive correlation = high test-retest reliability.
Example. A personality inventory given to 200 adults in January and again in March. If the January and March scores correlate at r = 0.90, the test is highly reliable.
Traits should be stable; states should not. High test-retest reliability is expected for stable traits (IQ, personality). For mood measures, you might expect lower test-retest correlation because mood itself changes.
Inter-Rater Reliability
Inter-rater reliability. Two or more raters measure the same phenomenon; their ratings are correlated.
Example. Two judges score the same 50 gymnastics routines. If their scores agree closely, inter-rater reliability is high.
This matters whenever human judgment is involved - diagnosis, interviews, behavioral coding, grading of essays. Low inter-rater reliability means the measurement depends on who is measuring, which is a serious problem.
Intra-Rater Reliability
Intra-rater reliability. The same rater scores the same materials on two occasions. Measures whether a rater is consistent with themselves over time.
Parallel-Forms Reliability (Split-Half Method)
Parallel-forms reliability. Two versions of a test, designed to measure the same construct, are given to the same people. If scores correlate strongly, the forms are reliable.
Split-half method. Divide the items of a single test in half (e.g., odd vs. even items) and correlate the two halves. High correlation = high internal consistency between halves.
Example. The SAT has multiple “forms” administered across test dates. Parallel-forms reliability tells you that a student’s score wouldn’t dramatically depend on which form they happened to take.
Internal Consistency
Internal consistency. All items on a test should measure the same construct. If they do, scores on individual items should correlate with each other and with the total.
Cronbach’s alpha (α). The standard measure of internal consistency. Ranges from 0 to 1.
- α > 0.9 = excellent
- α > 0.7 = acceptable for research
- α < 0.6 = questionable
A test with a single dimension should have high α. If α is low, the test may be measuring multiple unrelated things at once.
Reliability vs. Validity, Once More
| Reliability | Validity | Description |
|---|---|---|
| High | High | The test is good. |
| High | Low | The test consistently measures the wrong thing. A broken scale that always reads 10 lbs high. |
| Low | High | The test measures the right thing, noisily. Would need averaging many trials to see the true value. |
| Low | Low | The test is useless. |
Reliability is necessary but not sufficient for validity. A test cannot be valid without being reliable, but it can be reliable without being valid.