Validity
Imagine two cooking scenarios. In the first, you follow a recipe perfectly in a professional test kitchen with calibrated equipment - every measurement is exact, every variable is controlled. But you only used one brand of flour that nobody else can buy. Your recipe works flawlessly in that kitchen, but you have no idea if it works anywhere else. In the second scenario, you cook in a real home kitchen with whatever flour is available. Your results are messier, but they are much more representative of what an actual home cook would experience.
The first scenario has high internal validity but low external validity. The second has the opposite. Every study faces this tension, and the MCAT expects you to evaluate both.
Internal Validity
Internal validity asks: “Did the independent variable actually cause the change in the dependent variable?”
A study has high internal validity when you can confidently say the treatment - and nothing else - produced the observed effect. This requires removing confounders, using proper controls, randomizing participants, and blinding where possible.
Threats to internal validity include:
- Confounding variables: A third factor explains the result
- Selection bias: Groups were not equivalent at baseline
- Maturation: Participants change naturally over time (they grow older, more experienced, or heal on their own)
- History: External events during the study affect the outcome (e.g., a pandemic starts mid-trial)
- Attrition: Participants who drop out differ systematically from those who stay
- Testing effects: Taking a pretest changes performance on the posttest
External Validity
External validity asks: “Do the results generalize to other populations, settings, and times?”
A study has high external validity when its findings apply beyond the specific participants, location, and conditions of the study. A drug tested only on 20-year-old male college students may not work the same way in elderly women. A therapy tested in a university lab may not work in a rural clinic.
Threats to external validity include:
- Non-representative sample: Only studying one demographic, age group, or geographic region
- Artificial setting: Laboratory conditions that do not reflect real-world complexity
- Hawthorne effect: Participants behave differently because they know they are being observed (the results may not replicate when observation stops)
- Temporal factors: Results from the 1970s may not apply today
The Internal-External Validity Trade-off
There is a natural tension between the two. Tightly controlling an experiment (to boost internal validity) often makes it less realistic (reducing external validity). Studying participants in their natural environment (to boost external validity) introduces confounders (reducing internal validity).
Randomized controlled trials in clinical settings tend to have high internal validity but limited external validity (strict inclusion criteria, controlled conditions). Observational studies in the general population tend to have higher external validity but lower internal validity (no randomization, more confounders).
Construct Validity
Construct validity asks: “Does the measurement tool actually measure the abstract concept it claims to measure?”
This is especially relevant in psychology and social science, where many variables are abstract constructs (intelligence, depression, anxiety, self-esteem).
- An IQ test has high construct validity for measuring cognitive ability if it correlates with other accepted measures of intelligence and predicts outcomes that intelligence should predict (academic performance, problem-solving ability).
- A “creativity test” that only measures vocabulary has low construct validity for creativity because vocabulary is not the same thing as creativity.
Two subtypes you should recognize:
- Convergent validity: The measure correlates with other measures of the same construct (an anxiety questionnaire correlates with physiological stress markers)
- Discriminant validity: The measure does not correlate with measures of unrelated constructs (an anxiety questionnaire does not correlate with a math ability test)
Face Validity
Face validity asks: “Does the test look like it measures what it is supposed to measure?”
This is the weakest form of validity - it is just a subjective judgment. A math test that contains math problems has high face validity. A math test that contains only word puzzles has low face validity, even if it turns out to predict math ability surprisingly well.
Face validity matters for participant buy-in (people are more motivated to complete a test that looks relevant), but it does not guarantee that the test actually measures the right thing.
Content and Criterion Validity
Two more types sometimes appear on the MCAT:
Content validity: Does the test cover all aspects of the construct? A biology exam that only tests genetics has low content validity for “biology knowledge” because it ignores ecology, physiology, and cell biology.
Criterion validity: Does the test predict a relevant real-world outcome? A medical school admissions test has high criterion validity if students who score well also perform well in medical school. Criterion validity has two subtypes:
- Predictive validity: The test predicts future performance
- Concurrent validity: The test correlates with a current gold-standard measure