Reliability

Reliability

4 min read Updated Apr 19, 2026

Reliability is consistency. A test is reliable when repeated measurements give similar results - across administrations, across raters, across alternate forms, or across items. A reliable test can still be wrong (reliably measuring the wrong thing is possible), but an unreliable test is unusable regardless.

Test-Retest Reliability

Test-retest reliability. The same test is administered to the same subjects on two occasions, and the scores are correlated. High positive correlation = high test-retest reliability.

Example. A personality inventory given to 200 adults in January and again in March. If the January and March scores correlate at r = 0.90, the test is highly reliable.

Traits should be stable; states should not. High test-retest reliability is expected for stable traits (IQ, personality). For mood measures, you might expect lower test-retest correlation because mood itself changes.

Inter-Rater Reliability

Inter-rater reliability. Two or more raters measure the same phenomenon; their ratings are correlated.

Example. Two judges score the same 50 gymnastics routines. If their scores agree closely, inter-rater reliability is high.

This matters whenever human judgment is involved - diagnosis, interviews, behavioral coding, grading of essays. Low inter-rater reliability means the measurement depends on who is measuring, which is a serious problem.

Intra-Rater Reliability

Intra-rater reliability. The same rater scores the same materials on two occasions. Measures whether a rater is consistent with themselves over time.

Parallel-Forms Reliability (Split-Half Method)

Parallel-forms reliability. Two versions of a test, designed to measure the same construct, are given to the same people. If scores correlate strongly, the forms are reliable.

Split-half method. Divide the items of a single test in half (e.g., odd vs. even items) and correlate the two halves. High correlation = high internal consistency between halves.

Example. The SAT has multiple “forms” administered across test dates. Parallel-forms reliability tells you that a student’s score wouldn’t dramatically depend on which form they happened to take.

Internal Consistency

Internal consistency. All items on a test should measure the same construct. If they do, scores on individual items should correlate with each other and with the total.

Cronbach’s alpha (α). The standard measure of internal consistency. Ranges from 0 to 1.

  • α > 0.9 = excellent
  • α > 0.7 = acceptable for research
  • α < 0.6 = questionable

A test with a single dimension should have high α. If α is low, the test may be measuring multiple unrelated things at once.

Reliability vs. Validity, Once More

ReliabilityValidityDescription
HighHighThe test is good.
HighLowThe test consistently measures the wrong thing. A broken scale that always reads 10 lbs high.
LowHighThe test measures the right thing, noisily. Would need averaging many trials to see the true value.
LowLowThe test is useless.

Reliability is necessary but not sufficient for validity. A test cannot be valid without being reliable, but it can be reliable without being valid.

Distinguish test-retest from inter-rater reliability.
Click to reveal answer
Test-retest: same subjects take the same test at two times; scores are correlated across time. Inter-rater: different raters score the same subjects; scores are correlated across raters. Both measure consistency but on different dimensions.
Is reliability sufficient for validity?
Click to reveal answer
No - reliability is necessary but not sufficient. A test can be highly reliable (consistent) yet invalid (consistently measures the wrong thing), like a bathroom scale that always reads 10 lbs high. But an unreliable test cannot be valid, because noise limits what any measurement can capture.
What does Cronbach's alpha measure, and what does α = 0.9 tell you?
Click to reveal answer
Cronbach's alpha measures internal consistency - how well items on a test correlate with each other, indicating they measure the same construct. α = 0.9 means excellent internal consistency; the items hang together as a single scale.