Reliability

Reliability

7 min read Updated Mar 26, 2026

You step on your bathroom scale three mornings in a row. It reads 160, 160, 160. Great - the scale is consistent. But then you weigh yourself on a calibrated medical scale and it reads 155. Your bathroom scale was reliable (it gave the same answer every time) but not valid (the answer was wrong). Now imagine a different scale that reads 155, 162, 148. This one is neither reliable nor valid. Reliability is about consistency. Validity is about accuracy. You need both for good science, and the MCAT tests the difference.

What Is Reliability?

Reliability is the degree to which a measurement yields consistent, reproducible results. A reliable test gives you the same answer (or very close) every time you use it under the same conditions.

Think of reliability as precision. A reliable thermometer reads 98.6 degrees Fahrenheit every time you check a healthy person’s temperature. It may or may not be accurate (that is validity), but it is consistent.

Test-Retest Reliability

Test-retest reliability measures whether the same test produces the same results when given to the same people at two different time points.

How it works: Give a test at Time 1. Wait a period (days, weeks). Give the same test at Time 2. Correlate the two sets of scores. A high correlation (close to 1.0) means high test-retest reliability.

Example: A depression questionnaire administered today and again two weeks later should produce similar scores (assuming the person’s depression has not actually changed).

Limitation: If the construct being measured naturally fluctuates (mood, anxiety, energy level), low test-retest reliability might reflect genuine change rather than a flawed test.

Inter-Rater Reliability

Inter-rater reliability (also called inter-observer reliability) measures whether different observers produce the same ratings or scores when evaluating the same thing.

How it works: Two or more raters independently rate the same set of participants, essays, or behaviors. You calculate the agreement between raters. High agreement means high inter-rater reliability.

Example: Two psychiatrists independently interview the same 50 patients and assign diagnoses. If they agree on 45 out of 50 diagnoses, inter-rater reliability is high.

Why it matters: If a measurement depends on human judgment (grading essays, rating behavior, reading imaging), inter-rater reliability makes sure the results don’t depend on which specific person does the evaluating.

Common statistics for inter-rater reliability include Cohen’s kappa (for two raters with categorical data) and intraclass correlation coefficient (ICC) (for continuous data or more than two raters).

Internal Consistency

Internal consistency measures whether items within a single test that are supposed to measure the same construct actually produce similar results.

How it works: Look at the correlations among all items on a test. If the items are measuring the same underlying construct, they should correlate with each other.

Cronbach’s α (α) is the most common measure. It ranges from 0 to 1, where values above 0.7 are generally considered acceptable. A depression questionnaire with high α means that people who score high on one question about sadness also tend to score high on other questions about sadness.

Split-half reliability is a simpler version: divide the test items into two halves (odd-numbered vs. even-numbered questions) and correlate the scores from each half. High correlation means high internal consistency.

Reliability Summary Table

TypeQuestion It AnswersMethod
Test-retestSame results over time?Give same test twice, correlate scores
Inter-raterDifferent observers agree?Multiple raters score same items, measure agreement
Internal consistencyItems within test agree?Cronbach’s α or split-half correlation

Reliability vs. Validity: The Critical Relationship

Four bullseye targets illustrating the combinations of accuracy and precision: high accuracy and high precision (tight cluster at center), high accuracy but low precision (scattered around center), low accuracy but high precision (tight cluster off-center), and low accuracy and low precision (scattered off-center)
Accuracy (validity) vs. precision (reliability) illustrated with dartboard targets. Reliable but not valid: darts cluster tightly but miss the bullseye. Valid and reliable: darts cluster tightly around the bullseye. A measurement must be reliable (consistent) before it can be valid (accurate). Credit: Wikimedia Commons, CC BY-SA 3.0

This is one of the most tested concepts in research design on the MCAT:

Reliability is NECESSARY but NOT SUFFICIENT for validity.

A test must first be reliable (consistent) before it can be valid (accurate). If a test gives different results every time, it cannot possibly be measuring the right thing - it is not measuring anything consistently. But a test can be perfectly reliable and still not valid - your bathroom scale reads 160 every single time, but you actually weigh 155.

The four possible combinations:

ReliableNot Reliable
ValidIdeal: consistent and accurateImpossible: cannot be accurate if inconsistent
Not ValidConsistent but wrong (bathroom scale reads 160 every time, but true weight is 155)Inconsistent and wrong (worst case)
A personality test produces very different scores each time the same person takes it. Is this a problem with reliability, validity, or both?
Click to reveal answer
Both. The test has low test-retest reliability (inconsistent scores across time). Because reliability is necessary for validity, low reliability automatically means the test also lacks validity - it cannot be measuring the right thing if it cannot measure anything consistently.
Two radiologists independently read the same 100 MRI scans. They agree on the diagnosis in 95 out of 100 cases. What type of reliability is this, and is it high or low?
Click to reveal answer
This is inter-rater reliability, and it is high (95% agreement). When different observers produce the same results for the same data, the measurement is reliable across raters. This is especially important for diagnostic imaging, where the interpretation depends on human judgment.