Validity

Validity

6 min read Updated Apr 19, 2026

Validity asks: does this test actually measure what it claims to measure? The opposite, a test that is unreliable or measures the wrong thing, is at best useless and at worst misleading. MCAT passages regularly name a type of validity and expect you to know what it means.

Validity vs. Reliability

Before the types: do not confuse validity with reliability.

  • Validity. Accuracy. Is the test measuring the right construct?
  • Reliability. Consistency. Does the test give the same result on repeat measurement?

A broken bathroom scale that reads 10 pounds too high every time is reliable but not valid. A scale that gives you a different number each time, centered on your true weight, is valid on average but not reliable. You need both.

Internal Validity

Internal validity. The extent to which a study supports a causal conclusion within its own sample. High internal validity means that the observed effect is due to the manipulated variable, not to a confound.

Example. A drug trial is highly internally valid if randomization balances patient characteristics, dosing is consistent, and blinding prevents expectancy effects.

Threats to internal validity: confounding variables, selection bias, history effects, maturation, attrition.

External Validity

External validity. The extent to which the study’s findings generalize beyond the specific sample, setting, and conditions tested.

Example. A depression drug tested only on college-aged white men in one clinic may have strong internal validity but weak external validity because the sample is narrow.

Two subtypes are specifically named:

Population Validity

Population validity. Can the sample’s results extrapolate to the larger population? Requires a representative sample.

Example. A political poll with 200 respondents, all recruited from one Twitter account, has poor population validity for predicting nationwide voting.

Ecological Validity

Ecological validity. Do the study’s conditions resemble real-world settings? A laboratory memory task where people memorize word lists for money may measure something very different from how memory operates in everyday life.

Example. A driving safety study that uses a simulator has weaker ecological validity than one that uses actual cars in real traffic.

Construct Validity

Construct validity. The extent to which the test actually captures the abstract concept (the “construct”) it claims to measure.

Example. A test of “leadership” made entirely of vocabulary questions has poor construct validity - vocabulary is not what leadership is.

Two sub-types appear on the MCAT:

Convergent Validity

Convergent validity. The test correlates well with other established measures of the same construct.

Example. A new intelligence test should correlate strongly with the WAIS. If it doesn’t, it probably isn’t measuring intelligence.

Discriminant (Divergent) Validity

Discriminant validity. The test does not correlate with measures of unrelated constructs.

Example. A new depression inventory should NOT correlate strongly with a measure of extraversion. If it does, it may be confounding mood with personality.

Convergent and discriminant validity are paired: good construct validity means converging with similar measures AND diverging from unrelated ones.

Content Validity

Content validity. The extent to which the test covers the full range of the construct it claims to measure.

Example. A driver’s license exam that only tests parking has poor content validity, because driving also includes highway speeds, merging, and navigation. A content-valid exam samples across the whole domain.

Content validity is evaluated largely by expert judgment about domain coverage, not by statistical correlation.

Face Validity

Face validity. At a glance, does the test look like it measures what it claims to?

Example. A math test full of arithmetic problems has face validity for math ability. A math test that asks about your favorite color does not.

Face validity is the weakest form of validity - it is based on surface appearance only. A test can have high face validity and poor construct validity (e.g., a “stress questionnaire” that just asks about recent bad weather). But face validity matters for participant buy-in: if the test doesn’t look relevant, people may not take it seriously.

Criterion Validity

Criterion validity. The extent to which a test correlates with an outcome (“criterion”) it should predict.

Two sub-types, distinguished by when the criterion is measured:

Concurrent Validity

Concurrent validity. The test and the criterion are measured at the same time, and they correlate well.

Example. A new rapid depression screener is given to patients at the same time as the long-form BDI-II. If the two scores correlate strongly, the screener has concurrent validity.

Predictive Validity

Predictive validity. The test, given now, correlates with a criterion measured later.

Example. SAT scores taken in 12th grade correlate with college GPA measured 4 years later. The SAT has predictive validity for college performance.

Contrast concurrent and predictive. Concurrent = same-time benchmark. Predictive = future outcome. Both are types of criterion validity.

TypeQuestion It AnswersExample
InternalIs the effect in THIS study causal?RCT with randomization
ExternalWill the finding generalize?Diverse, representative sample
PopulationTo the whole population?National probability sample
EcologicalTo the real world?Field vs. lab
ConstructDoes it measure the abstract concept?Leadership test correlates with leadership behavior
ConvergentMatch with similar measures?New IQ test ≈ WAIS
DiscriminantDiffer from unrelated measures?Depression ≠ extraversion
ContentCover the whole domain?Driving test covers all skills
FaceLook right at a glance?Math test has math problems
CriterionCorrelate with an outcome?SAT and college grades
ConcurrentSame-time benchmark?Quick screener vs. full inventory
PredictiveFuture outcome?SAT predicts GPA
Distinguish convergent from discriminant validity with a simple example.
Click to reveal answer
Convergent: a new depression test correlates strongly with an established depression inventory (similar constructs should agree). Discriminant: the same test does NOT correlate with a measure of extraversion (unrelated constructs should be separable). Good construct validity requires BOTH.
What is ecological validity?
Click to reveal answer
Ecological validity is whether the study's conditions and tasks resemble real-world situations enough for findings to apply outside the lab. A driving study done in a simulator has weaker ecological validity than one in real traffic. It is a type of external validity.
Distinguish concurrent from predictive validity.
Click to reveal answer
Both are types of criterion validity. Concurrent validity: the test and the criterion are measured at the same time (e.g., brief screener vs. full inventory today). Predictive validity: the test is measured now and correlates with a future criterion (e.g., SAT today predicts college GPA 4 years later).
Why is face validity the weakest form of validity?
Click to reveal answer
Face validity is based only on surface appearance - whether the test looks right at a glance. A test can look highly relevant but fail to measure the construct (poor construct validity) or miss key parts of it (poor content validity). Face validity matters for participant compliance but is not evidence the test works.