Validity
Validity asks: does this test actually measure what it claims to measure? The opposite, a test that is unreliable or measures the wrong thing, is at best useless and at worst misleading. MCAT passages regularly name a type of validity and expect you to know what it means.
Validity vs. Reliability
Before the types: do not confuse validity with reliability.
- Validity. Accuracy. Is the test measuring the right construct?
- Reliability. Consistency. Does the test give the same result on repeat measurement?
A broken bathroom scale that reads 10 pounds too high every time is reliable but not valid. A scale that gives you a different number each time, centered on your true weight, is valid on average but not reliable. You need both.
Internal Validity
Internal validity. The extent to which a study supports a causal conclusion within its own sample. High internal validity means that the observed effect is due to the manipulated variable, not to a confound.
Example. A drug trial is highly internally valid if randomization balances patient characteristics, dosing is consistent, and blinding prevents expectancy effects.
Threats to internal validity: confounding variables, selection bias, history effects, maturation, attrition.
External Validity
External validity. The extent to which the study’s findings generalize beyond the specific sample, setting, and conditions tested.
Example. A depression drug tested only on college-aged white men in one clinic may have strong internal validity but weak external validity because the sample is narrow.
Two subtypes are specifically named:
Population Validity
Population validity. Can the sample’s results extrapolate to the larger population? Requires a representative sample.
Example. A political poll with 200 respondents, all recruited from one Twitter account, has poor population validity for predicting nationwide voting.
Ecological Validity
Ecological validity. Do the study’s conditions resemble real-world settings? A laboratory memory task where people memorize word lists for money may measure something very different from how memory operates in everyday life.
Example. A driving safety study that uses a simulator has weaker ecological validity than one that uses actual cars in real traffic.
Construct Validity
Construct validity. The extent to which the test actually captures the abstract concept (the “construct”) it claims to measure.
Example. A test of “leadership” made entirely of vocabulary questions has poor construct validity - vocabulary is not what leadership is.
Two sub-types appear on the MCAT:
Convergent Validity
Convergent validity. The test correlates well with other established measures of the same construct.
Example. A new intelligence test should correlate strongly with the WAIS. If it doesn’t, it probably isn’t measuring intelligence.
Discriminant (Divergent) Validity
Discriminant validity. The test does not correlate with measures of unrelated constructs.
Example. A new depression inventory should NOT correlate strongly with a measure of extraversion. If it does, it may be confounding mood with personality.
Convergent and discriminant validity are paired: good construct validity means converging with similar measures AND diverging from unrelated ones.
Content Validity
Content validity. The extent to which the test covers the full range of the construct it claims to measure.
Example. A driver’s license exam that only tests parking has poor content validity, because driving also includes highway speeds, merging, and navigation. A content-valid exam samples across the whole domain.
Content validity is evaluated largely by expert judgment about domain coverage, not by statistical correlation.
Face Validity
Face validity. At a glance, does the test look like it measures what it claims to?
Example. A math test full of arithmetic problems has face validity for math ability. A math test that asks about your favorite color does not.
Face validity is the weakest form of validity - it is based on surface appearance only. A test can have high face validity and poor construct validity (e.g., a “stress questionnaire” that just asks about recent bad weather). But face validity matters for participant buy-in: if the test doesn’t look relevant, people may not take it seriously.
Criterion Validity
Criterion validity. The extent to which a test correlates with an outcome (“criterion”) it should predict.
Two sub-types, distinguished by when the criterion is measured:
Concurrent Validity
Concurrent validity. The test and the criterion are measured at the same time, and they correlate well.
Example. A new rapid depression screener is given to patients at the same time as the long-form BDI-II. If the two scores correlate strongly, the screener has concurrent validity.
Predictive Validity
Predictive validity. The test, given now, correlates with a criterion measured later.
Example. SAT scores taken in 12th grade correlate with college GPA measured 4 years later. The SAT has predictive validity for college performance.
Contrast concurrent and predictive. Concurrent = same-time benchmark. Predictive = future outcome. Both are types of criterion validity.
| Type | Question It Answers | Example |
|---|---|---|
| Internal | Is the effect in THIS study causal? | RCT with randomization |
| External | Will the finding generalize? | Diverse, representative sample |
| Population | To the whole population? | National probability sample |
| Ecological | To the real world? | Field vs. lab |
| Construct | Does it measure the abstract concept? | Leadership test correlates with leadership behavior |
| Convergent | Match with similar measures? | New IQ test ≈ WAIS |
| Discriminant | Differ from unrelated measures? | Depression ≠ extraversion |
| Content | Cover the whole domain? | Driving test covers all skills |
| Face | Look right at a glance? | Math test has math problems |
| Criterion | Correlate with an outcome? | SAT and college grades |
| Concurrent | Same-time benchmark? | Quick screener vs. full inventory |
| Predictive | Future outcome? | SAT predicts GPA |