Reliability
You step on your bathroom scale three mornings in a row. It reads 160, 160, 160. Great - the scale is consistent. But then you weigh yourself on a calibrated medical scale and it reads 155. Your bathroom scale was reliable (it gave the same answer every time) but not valid (the answer was wrong). Now imagine a different scale that reads 155, 162, 148. This one is neither reliable nor valid. Reliability is about consistency. Validity is about accuracy. You need both for good science, and the MCAT tests the difference.
What Is Reliability?
Reliability is the degree to which a measurement yields consistent, reproducible results. A reliable test gives you the same answer (or very close) every time you use it under the same conditions.
Think of reliability as precision. A reliable thermometer reads 98.6 degrees Fahrenheit every time you check a healthy person’s temperature. It may or may not be accurate (that is validity), but it is consistent.
Test-Retest Reliability
Test-retest reliability measures whether the same test produces the same results when given to the same people at two different time points.
How it works: Give a test at Time 1. Wait a period (days, weeks). Give the same test at Time 2. Correlate the two sets of scores. A high correlation (close to 1.0) means high test-retest reliability.
Example: A depression questionnaire administered today and again two weeks later should produce similar scores (assuming the person’s depression has not actually changed).
Limitation: If the construct being measured naturally fluctuates (mood, anxiety, energy level), low test-retest reliability might reflect genuine change rather than a flawed test.
Inter-Rater Reliability
Inter-rater reliability (also called inter-observer reliability) measures whether different observers produce the same ratings or scores when evaluating the same thing.
How it works: Two or more raters independently rate the same set of participants, essays, or behaviors. You calculate the agreement between raters. High agreement means high inter-rater reliability.
Example: Two psychiatrists independently interview the same 50 patients and assign diagnoses. If they agree on 45 out of 50 diagnoses, inter-rater reliability is high.
Why it matters: If a measurement depends on human judgment (grading essays, rating behavior, reading imaging), inter-rater reliability makes sure the results don’t depend on which specific person does the evaluating.
Common statistics for inter-rater reliability include Cohen’s kappa (for two raters with categorical data) and intraclass correlation coefficient (ICC) (for continuous data or more than two raters).
Internal Consistency
Internal consistency measures whether items within a single test that are supposed to measure the same construct actually produce similar results.
How it works: Look at the correlations among all items on a test. If the items are measuring the same underlying construct, they should correlate with each other.
Cronbach’s α (α) is the most common measure. It ranges from 0 to 1, where values above 0.7 are generally considered acceptable. A depression questionnaire with high α means that people who score high on one question about sadness also tend to score high on other questions about sadness.
Split-half reliability is a simpler version: divide the test items into two halves (odd-numbered vs. even-numbered questions) and correlate the scores from each half. High correlation means high internal consistency.
Reliability Summary Table
| Type | Question It Answers | Method |
|---|---|---|
| Test-retest | Same results over time? | Give same test twice, correlate scores |
| Inter-rater | Different observers agree? | Multiple raters score same items, measure agreement |
| Internal consistency | Items within test agree? | Cronbach’s α or split-half correlation |
Reliability vs. Validity: The Critical Relationship
This is one of the most tested concepts in research design on the MCAT:
Reliability is NECESSARY but NOT SUFFICIENT for validity.
A test must first be reliable (consistent) before it can be valid (accurate). If a test gives different results every time, it cannot possibly be measuring the right thing - it is not measuring anything consistently. But a test can be perfectly reliable and still not valid - your bathroom scale reads 160 every single time, but you actually weigh 155.
The four possible combinations:
| Reliable | Not Reliable | |
|---|---|---|
| Valid | Ideal: consistent and accurate | Impossible: cannot be accurate if inconsistent |
| Not Valid | Consistent but wrong (bathroom scale reads 160 every time, but true weight is 155) | Inconsistent and wrong (worst case) |