Imagine you’re a detective arriving at a crime scene. You look around and gather clues (observation). You notice the window is broken from the outside: how did the intruder get in? (question). You form a theory about what happened (hypothesis). You dust for fingerprints and check security footage (experiment). You compare the evidence to your theory (analysis). Either the evidence supports your theory or it doesn’t (conclusion).
Science works exactly the same way. Every experiment is a detective story. The scientific method isn’t a rigid checklist — it’s a cycle of asking questions, testing ideas, and revising your understanding based on evidence. The MCAT expects you to recognize these steps, identify where a study sits in the cycle, and judge whether the conclusions follow from the data.
The Steps
The scientific method as an iterative cycle. Observations lead to questions, which generate testable hypotheses. Experiments test those hypotheses, and the results either support the hypothesis or prompt revision and further investigation. The process is continuous and self-correcting. Credit: Wikimedia Commons, CC BY-SA 3.0
The scientific method follows a general sequence, though in practice scientists often jump between steps or go back to earlier ones.
1. Observation. Something catches your attention. A patient with a rare disease improves after taking an unrelated medication. Cells in a petri dish grow faster at a certain temperature. Data from a survey shows an unexpected pattern.
2. Question. You ask why or how. Why did that patient improve? What is it about that temperature that accelerates cell growth?
3. Hypothesis. You propose a testable, falsifiable explanation. “The medication reduces inflammation by blocking receptor X.” This is not a guess - it is a specific, mechanistic prediction that can be proven wrong.
4. Experiment. You design a study to test the hypothesis. This is where variables, controls, and blinding come in (we will cover those in the next several sections).
5. Data collection and analysis. You gather results and apply statistical methods to see whether the data support or contradict the hypothesis.
6. Conclusion. You interpret the results. Did the data support the hypothesis? If yes, the hypothesis survives (but is not “proven” - more on that below). If no, you revise the hypothesis and start again.
Hypotheses Must Be Testable and Falsifiable
A hypothesis is only scientific if it can, in principle, be shown to be wrong. This is the criterion of falsifiability, introduced by philosopher Karl Popper.
Testable: You can design an experiment to evaluate it. “Drug X lowers blood pressure” is testable - give the drug to patients and measure their blood pressure.
Falsifiable: There is a possible outcome that would disprove it. If blood pressure does not drop, the hypothesis is falsified.
A statement like “everything happens for a reason” is not falsifiable because no observation could ever disprove it. It may be a perfectly fine philosophical idea, but it is not a scientific hypothesis.
Hypothesis vs. Theory vs. Law
These three terms are not a hierarchy of certainty. They describe different things entirely.
Term
What It Is
Example
Hypothesis
A specific, testable prediction about a single phenomenon
”This drug lowers blood pressure by blocking ACE”
Theory
A broad, well-tested explanation supported by a large body of evidence
Theory of evolution, germ theory of disease
Law
A concise mathematical description of a consistent relationship in nature
Newton’s second law (F = ma), Boyle’s law (PV = constant)
A theory does not “graduate” into a law. Laws describe what happens (the pattern). Theories explain why it happens (the mechanism). Both are supported by extensive evidence, but they serve different purposes.
Inductive vs. Deductive Reasoning
Scientists use two complementary reasoning approaches.
Inductive reasoning moves from specific observations to a general conclusion. You observe that every swan you have ever seen is white, so you conclude “all swans are white.” Inductive conclusions are probable but never certain - a single black swan disproves the rule.
Deductive reasoning moves from a general principle to a specific prediction. If all mammals produce milk, and a whale is a mammal, then a whale produces milk. Deductive conclusions are certain if the premises are true.
The scientific method uses both: inductive reasoning to form hypotheses from observations, and deductive reasoning to generate testable predictions from those hypotheses.
Null and Alternative Hypotheses
In formal research, hypotheses come in pairs.
The null hypothesis (H0) states that there is no effect or no difference. “The drug has no effect on blood pressure compared to placebo.”
The alternative hypothesis (H1 or Hₐ) states that there is an effect or a difference. “The drug lowers blood pressure compared to placebo.”
Experiments are designed to reject the null hypothesis. If the data show a statistically significant difference, you reject H0 in favor of H1. If they do not, you fail to reject H0. Note the language: you never “accept” the null hypothesis - you simply fail to reject it, because absence of evidence is not evidence of absence.
What is the difference between a theory and a law in science?
Click to reveal answer
A law is a concise mathematical description of a consistent natural pattern (describes what happens). A theory is a broad, well-tested explanation of why a phenomenon occurs (explains why it happens). They are not a hierarchy - a theory does not become a law. Both are supported by extensive evidence.
A researcher proposes that "negative energy in a room causes illness." Is this a valid scientific hypothesis? Why or why not?
Click to reveal answer
No. This is not a valid scientific hypothesis because it is not falsifiable. "Negative energy" is not operationally defined, cannot be measured, and no experiment could disprove the claim. A scientific hypothesis must make a specific, measurable prediction that could be shown to be wrong.
Think about baking cookies. You want to test whether more sugar makes cookies taste better. The amount of sugar you add is the thing you deliberately change — that’s your independent variable. The taste rating your friends give each batch is the thing you measure — that’s your dependent variable.
But what if you also accidentally used a different oven temperature for each batch? Now you can’t tell whether the taste difference came from the sugar or from the temperature. That uninvited guest is a confounding variable, and it can quietly ruin an otherwise good experiment.
Every experiment boils down to this: change one thing, measure another thing, and make sure nothing else sneaks in to muddy the results. The MCAT tests this concept relentlessly. If you can identify the independent, dependent, and confounding variables in a passage, you can answer most research design questions.
The Three Core Variables
Independent variable (IV): The variable the researcher deliberately manipulates or changes. In a drug trial, the IV is the drug dose. In a psychology study, the IV might be the type of therapy given.
Dependent variable (DV): The variable the researcher measures as an outcome. It “depends” on the IV. In a drug trial, the DV might be blood pressure. In a psychology study, the DV might be depression score.
Confounding variable: Any variable other than the IV that could influence the DV and was not properly controlled. Confounders threaten the validity of the experiment because they provide an alternative explanation for the results.
Graphing Convention
When plotting experimental results, there is a universal convention:
Independent variable (IV) goes on the x-axis (horizontal)
Dependent variable (DV) goes on the y-axis (vertical)
This is not arbitrary. The x-axis represents what the researcher controls, and the y-axis represents what responds. If an MCAT passage shows a graph, you can immediately identify the IV and DV by looking at the axis labels.
Extraneous vs. Confounding Variables
A confounding variable (Z) is associated with both the independent variable (X) and the dependent variable (Y), creating a spurious apparent relationship between X and Y. Without controlling for Z, a researcher might incorrectly conclude that X causes Y. Credit: Wikimedia Commons, CC BY-SA 3.0
These two terms are related but not identical.
Extraneous variables are any variables other than the IV that could affect the DV. Room temperature, time of day, participant mood - all are extraneous in most studies.
Confounding variables are extraneous variables that actually do vary systematically with the IV, making it impossible to separate their effect from the IV’s. Not every extraneous variable becomes a confounder - only the ones that correlate with the treatment.
Example: A researcher tests whether a new study method improves test scores. The experimental group uses the new method and studies in the morning. The control group uses the old method and studies at night. Time of day is extraneous, but because it varies systematically with the study method (all morning students got the new method), it becomes a confounder.
Controlled (Constant) Variables
Controlled variables are factors the researcher deliberately keeps the same across all groups. If you are testing a drug’s effect on blood pressure, you might control for age, sex, diet, and exercise by matching the groups or holding those factors constant.
Controlling variables is how researchers prevent extraneous variables from becoming confounders. The more variables you control, the more confident you can be that the IV caused the change in the DV.
Operationalization
Operationalization is the process of defining an abstract concept in terms of specific, measurable procedures. This is essential because many variables - especially in psychology and social science - are abstract.
“Intelligence” is abstract. “Score on a standardized IQ test” is operationalized.
“Stress” is abstract. “Cortisol level in saliva” is operationalized.
“Aggression” is abstract. “Number of times a participant presses a button to deliver a loud noise to another person” is operationalized.
Levels of the Independent Variable
The IV often has multiple levels (also called conditions or groups). A drug trial might have three levels: placebo, low dose, and high dose. Each level is a different value of the IV that participants are assigned to.
When there are only two levels, you typically have an experimental group (receives treatment) and a control group (does not). When there are more than two levels, the study can reveal dose-response relationships or compare multiple treatments.
A researcher gives Group A a new drug and Group B a placebo, then measures blood glucose levels. Identify the IV and DV.
Click to reveal answer
IV: drug condition (new drug vs. placebo) - this is what the researcher manipulates. DV: blood glucose level - this is what the researcher measures as an outcome. The IV goes on the x-axis, and the DV goes on the y-axis.
In a study on exercise and depression, participants who exercise also happen to have higher income. Why is income a confounding variable?
Click to reveal answer
Income is a confounding variable because it varies systematically with the IV (exercise) and could independently affect the DV (depression). Higher income may reduce depression through better healthcare, less financial stress, or other mechanisms. The researcher cannot tell whether lower depression scores are due to exercise or to higher income.
Suppose you buy a new pregnancy test and want to make sure it actually works before relying on the result. You would run it on a sample you know is positive (to confirm the test can detect pregnancy) and a sample you know is negative (to confirm the test does not give false alarms). Those two checks are controls. Without them, you have no idea whether the test works at all - a positive result might mean the test shows positive for everything, and a negative result might mean it is broken.
Controls are the backbone of any experiment. They are the reference points that let you interpret your results. The MCAT loves to describe an experiment and then ask which control is missing, which control failed, or what the absence of a control means for the conclusions. Get these straight and those questions become free points.
Positive Control
A positive control is a condition that is known to produce the expected effect. Its purpose is to confirm that the experimental setup is working properly.
If you are testing whether a new antibiotic kills bacteria, your positive control is a plate treated with a known effective antibiotic. If the bacteria on the positive control plate do not die, something is wrong with your experimental procedure - maybe the bacteria are resistant, or the incubation conditions are off. The positive control catches equipment or procedural failures.
Negative Control
A negative control is a condition that is known to produce no effect. Its purpose is to confirm that the experiment is not generating false positives.
In the antibiotic experiment, your negative control is a plate with no antibiotic at all (just the growth medium). If bacteria on the negative control plate die anyway, something in the environment is killing them - contamination, wrong temperature, or a problem with the medium. The negative control catches false positives.
Quick Reference
Control Type
What It Shows
If It Fails
Positive control
The experiment can detect an effect
The setup is broken - you cannot trust negative results
Negative control
The experiment does not produce false effects
Something besides your IV is causing the effect - you cannot trust positive results
Placebo and Sham Controls
A placebo is an inactive treatment designed to look identical to the real treatment. In a drug trial, the placebo is typically a sugar pill that matches the real pill in size, shape, and color. The placebo group accounts for the psychological effect of believing you are being treated (the placebo effect).
A sham procedure is the surgical equivalent of a placebo. If the experimental group receives a new knee surgery, the sham group might get an incision and stitches without the actual surgical repair. Sham controls are controversial but powerful because they isolate the effect of the procedure itself from the effect of the overall surgical experience.
A vehicle control is used when the treatment is dissolved in a solvent (the “vehicle”). The vehicle control group receives the solvent alone, without the active compound. This ensures that any effect is due to the drug, not the solvent.
Control Group vs. Experimental Group
The control group does not receive the treatment (or receives a placebo). The experimental group receives the actual treatment being tested. By comparing outcomes between the two groups, the researcher can tell whether the treatment had an effect.
A well-designed experiment keeps everything identical between the control and experimental groups except for the independent variable. Same age range, same measurement procedures, same environmental conditions. The only difference should be the treatment itself.
Why Controls Matter on the MCAT
The most common MCAT question pattern for controls is: “A researcher observes Effect X in the experimental group. Which of the following is a valid criticism of this study?” The answer usually involves a missing or inadequate control. Specifically:
If there is no negative control, you cannot rule out that the effect happens without the treatment
If there is no positive control, you cannot confirm the experiment could detect an effect
If the placebo does not match the treatment in appearance, blinding breaks down
In a drug trial, the negative control group also shows significant improvement. What does this mean for the study's conclusions?
Click to reveal answer
The study's conclusions are compromised. If the negative control (no treatment) shows improvement, something other than the drug is causing the effect - possibly spontaneous recovery, placebo effect from attention, or a flaw in the experimental setup. You cannot attribute improvement in the experimental group to the drug alone.
A researcher tests a new antibiotic on bacterial cultures. The positive control (known antibiotic) fails to kill the bacteria. What should the researcher conclude?
Click to reveal answer
The experimental setup is not working properly. Since the positive control (a known effective antibiotic) failed, the researcher cannot trust any results from this experiment. The bacteria may be resistant, the culture conditions may be wrong, or the antibiotic may have degraded. The experiment must be repeated with a functional positive control before drawing conclusions.
You want to test whether coffee improves reaction time. You have two options. Option one: split 100 people into two groups - give Group A coffee and Group B water, then compare their reaction times. Option two: give all 100 people water on Monday and coffee on Tuesday, then compare each person’s reaction time across the two days. These two approaches represent the two fundamental experimental designs, and each has trade-offs that the MCAT expects you to know.
Between-Subjects Design
In a between-subjects design, different participants are assigned to each condition. Group A gets the treatment, Group B gets the control, and you compare the two groups.
Advantage: No order effects. Participants only experience one condition, so fatigue, practice, or boredom from the first condition cannot contaminate the second.
Disadvantage: Individual differences. People in Group A might naturally have faster reaction times than people in Group B, regardless of coffee. You need larger sample sizes to overcome this variability, and you must use random assignment to spread individual differences evenly.
Within-Subjects Design (Repeated Measures)
In a within-subjects design, the same participants experience every condition. Each person serves as their own control. You measure the DV under condition A, then again under condition B, and compare.
Advantage: Eliminates individual differences. Since each person is compared to themselves, natural variability between people is removed. You need fewer participants.
Disadvantage: Order effects. Going through one condition first can affect performance in the second condition. If everyone drinks coffee first and water second, maybe their reaction time on the water day is worse because they are tired from the previous session, not because water is inferior to coffee.
Counterbalancing
Counterbalancing is the solution to order effects in within-subjects designs. You vary the sequence of conditions across participants.
Half the participants get coffee first and water second. The other half get water first and coffee second. Any order effect (practice, fatigue, carryover) is spread equally across both conditions, so it cannot systematically favor one over the other.
A crossover design is a specific type of within-subjects design where participants switch between treatment and control, usually with a “washout period” in between to let the first treatment’s effects fade.
Quick Comparison
Feature
Between-Subjects
Within-Subjects
Participants per condition
Different people
Same people
Individual differences
A concern (need random assignment)
Eliminated (each person is own control)
Order effects
Not a concern
A concern (need counterbalancing)
Sample size needed
Larger
Smaller
Also called
Independent groups design
Repeated measures design
Random Assignment vs. Random Selection
These two terms sound similar but address completely different problems.
Random assignment means every participant has an equal chance of being placed in any condition (treatment or control). It is used to distribute confounding variables evenly across groups. Random assignment supports internal validity - the confidence that the IV caused the change in the DV.
Random selection (random sampling) means every member of the population has an equal chance of being chosen for the study. It is used to make sure the sample represents the broader population. Random selection supports external validity - the confidence that results generalize beyond the study.
Factorial Designs
A factorial design tests two or more independent variables at the same time. Instead of just testing Drug vs. Placebo, you might test Drug vs. Placebo AND Therapy vs. No Therapy. This creates a 2 x 2 design with four conditions:
Therapy
No Therapy
Drug
Drug + Therapy
Drug + No Therapy
Placebo
Placebo + Therapy
Placebo + No Therapy
The power of factorial designs is that they reveal interaction effects - situations where the effect of one IV depends on the level of the other IV. For example, the drug might work only when combined with therapy but not alone.
Matched-Pairs Design
A matched-pairs design is a hybrid approach. Participants are paired based on a relevant characteristic (age, severity of illness, baseline score), and then one member of each pair is randomly assigned to the treatment and the other to the control.
This reduces individual differences (like within-subjects) while avoiding order effects (like between-subjects). It is especially useful when the characteristic being matched is known to strongly affect the DV.
A study uses the same 30 participants to test three different study methods. What design is this, and what is the biggest threat to validity?
Click to reveal answer
This is a within-subjects (repeated measures) design. The biggest threat is order effects - participants may perform better on later methods due to practice, or worse due to fatigue. The researcher should use counterbalancing (varying the order of methods across participants) to address this.
What is the difference between random assignment and random selection?
Click to reveal answer
Random assignment: every participant has an equal chance of being placed in any experimental condition (treatment or control). Supports internal validity. Random selection: every member of the target population has an equal chance of being chosen for the study. Supports external validity. A study can use one without the other.
Picture a taste test at a grocery store. A company wants to know whether customers prefer their cola over the competitor’s. If the cups are labeled “Our Brand” and “Other Brand,” most people will say they prefer “Our Brand” - they are biased by the label, not the flavor. But if both cups are unlabeled and participants do not know which is which, the test measures actual taste preference. That is blinding in action: removing information that could bias the result.
Randomization and blinding are the two most powerful tools for eliminating bias in experiments. Randomization makes sure groups are comparable before the study starts. Blinding makes sure expectations don’t distort the results once the study is underway. Together, they are what separate a rigorous experiment from a glorified anecdote.
Randomization
Randomization means assigning participants to groups using a chance process, so every participant has an equal probability of ending up in any group.
Why does this matter? Because humans are not identical. They differ in age, genetics, health, motivation, and countless other factors. If the researcher picks who goes into which group (or if participants self-select), those differences may cluster in one group and create confounders. Randomization distributes these known and unknown confounders roughly equally across groups.
Common randomization methods include:
Simple randomization: Coin flip or random number generator for each participant. Easy but can produce unequal group sizes, especially with small samples.
Block randomization: Participants are randomized in blocks (e.g., groups of 4) to keep group sizes equal at regular intervals.
Stratified randomization: Participants are first sorted by a key characteristic (e.g., sex or age), then randomized within each stratum. This guarantees balance on that characteristic.
Single-Blind Design
In a single-blind study, the participant does not know which group they are in (treatment or control), but the researcher does know.
This eliminates the placebo effect - the tendency for participants to feel better simply because they believe they are receiving treatment. If participants know they are in the control group, they may not try as hard, report less improvement, or drop out.
However, single-blind designs still leave room for observer bias (also called experimenter bias). If the researcher knows which participants received the drug, they might unconsciously evaluate those participants more favorably, measure their outcomes differently, or prompt them with leading questions.
Double-Blind Design
In a double-blind study, neither the participant nor the researcher knows which group the participant is in. A third party (such as a pharmacist) assigns the treatments using a code that is not revealed until the study is complete.
Double-blinding eliminates both the placebo effect and observer bias. It is the standard for clinical drug trials and is considered one of the hallmarks of rigorous experimental design.
Triple-Blind Design
In a triple-blind study, the participant, the researcher, and the data analyst do not know group assignments. The analyst receives coded data and performs statistical tests without knowing which code corresponds to treatment vs. control.
This prevents the analyst from (consciously or unconsciously) choosing statistical methods or interpretations that favor a particular outcome.
Summary Table
Design
Who Is Blinded
What It Prevents
Open-label (unblinded)
Nobody
Nothing - highest risk of bias
Single-blind
Participant
Placebo effect
Double-blind
Participant + researcher
Placebo effect + observer bias
Triple-blind
Participant + researcher + analyst
Placebo effect + observer bias + analysis bias
When Blinding Is Not Possible
Some studies cannot be blinded. If the IV is a type of surgery, the surgeon obviously knows which procedure they performed. If the IV is exercise vs. no exercise, participants know whether they are exercising. In these cases, researchers use other strategies to minimize bias:
Blinded outcome assessors: Even if the participant and treating physician know the group, the person measuring the outcome (e.g., a radiologist reading a scan) does not.
Objective outcome measures: Using lab values or imaging instead of patient-reported symptoms reduces the influence of expectations.
Sham procedures: A fake surgery that mimics the experience without the actual intervention.
Why is a double-blind design considered superior to a single-blind design?
Click to reveal answer
A double-blind design eliminates both the placebo effect (participant does not know their group) and observer bias (researcher does not know which participants received treatment). A single-blind design only removes one source of bias - the participant's expectations - but the researcher can still unconsciously influence outcomes.
A researcher randomizes 20 patients into treatment and control groups, but by chance, 8 of the 10 oldest patients end up in the treatment group. Is randomization "broken"?
Click to reveal answer
Randomization is not broken, but it may have failed to balance confounders in this small sample. With only 20 participants, chance imbalances are common. This is why larger sample sizes are important - the law of large numbers makes randomization more effective at distributing confounders evenly. The researcher could also use stratified randomization by age to prevent this.
Not every research question can be answered with an experiment. You cannot randomly assign people to smoke for 20 years to see if it causes cancer - that would be unethical. Instead, you observe people who already smoke and compare them to people who do not. This is the world of observational studies: the researcher watches and measures but does not manipulate anything.
The critical trade-off is this: because the researcher does not control the independent variable, observational studies cannot establish causation. They can only identify associations. The MCAT tests this distinction heavily and expects you to know the three major types of observational studies, what each can and cannot tell you, and how to identify them from a passage description.
Cross-Sectional Studies
A cross-sectional study collects data from a group of people at a single point in time. Think of it as a photograph - you capture one moment and examine the patterns.
What it measures: Prevalence (how common a condition is at that moment). For example: “What percentage of college students currently experience anxiety?”
What it cannot do: Establish temporal sequence. Because everything is measured at the same time, you can’t tell which came first. If people with anxiety also exercise less, did anxiety cause them to stop exercising, or did lack of exercise cause anxiety?
Strengths: Fast, cheap, no follow-up needed. Good for measuring prevalence and generating hypotheses.
Weaknesses: Cannot show causation or temporal sequence. Subject to prevalence-incidence bias (people with mild, short-lived conditions are underrepresented).
Case-Control Studies
A case-control study starts with the outcome and looks backward. The researcher identifies people who have the condition (cases) and people who do not (controls), then looks back to see whether the groups differed in their past exposures.
Direction: Retrospective (backward-looking). You start at the effect and trace back to possible causes.
What it measures: Odds ratio (OR) - the odds of exposure among cases compared to the odds of exposure among controls. It cannot directly measure incidence or relative risk because you started by selecting on the outcome, not the exposure.
Example: You identify 100 patients with lung cancer (cases) and 100 patients without lung cancer (controls). You look at their smoking history. If 80% of cases smoked but only 20% of controls smoked, smoking is strongly associated with lung cancer.
Strengths: Excellent for studying rare diseases (because you recruit cases directly). Fairly fast and inexpensive since the outcome has already occurred.
Weaknesses: Prone to recall bias (participants may not accurately remember past exposures). Cannot show causation or calculate incidence.
Cohort Studies
A cohort study starts with the exposure and follows forward. The researcher identifies a group of people who are exposed (e.g., smokers) and a group who are not exposed (e.g., non-smokers), then follows both groups over time to see who develops the outcome.
Direction: Prospective (forward-looking). You start at the cause and watch for the effect.
What it measures: Incidence (how many new cases develop over time), relative risk (RR), and attributable risk. Because you follow people from exposure to outcome, you can directly calculate how much more likely the exposed group is to develop the condition.
Example: You recruit 10,000 smokers and 10,000 non-smokers, then follow them for 20 years. You count how many in each group develop lung cancer. If 15% of smokers develop cancer vs. 1% of non-smokers, the relative risk is 15.
Strengths: Can establish temporal sequence (exposure came before outcome). Can calculate incidence and relative risk. Less susceptible to recall bias because data is collected as events happen.
Weaknesses: Expensive, slow (especially for diseases with long latency periods). Subject to attrition bias (participants may drop out over time).
Comparison Table
Feature
Cross-Sectional
Case-Control
Cohort
Time direction
Snapshot (one point)
Retrospective (backward)
Prospective (forward)
Starts with
Neither
Outcome (disease)
Exposure (risk factor)
Measures
Prevalence
Odds ratio
Incidence, relative risk
Can show causation?
No
No
Stronger, but still not definitive
Best for
Prevalence data, hypothesis generation
Rare diseases
Establishing temporal sequence
Cost/time
Lowest
Moderate
Highest
Decision Tree for Identifying Study Type
When an MCAT passage describes a study, ask three questions:
Did the researcher manipulate anything? If yes, it is an experiment. If no, it is observational.
Was data collected at one time or over time? If one time point, it is cross-sectional.
Did the researcher start with the outcome or the exposure? If they started by finding people with the disease and looking back, it is case-control. If they started with exposed vs. unexposed groups and followed forward, it is cohort.
Retrospective Cohort Studies
A retrospective cohort study is a hybrid. The researcher uses existing records (medical charts, employment databases) to identify a cohort that was exposed vs. unexposed in the past, then looks at outcomes that have already occurred. It follows the logic of a cohort study (exposure to outcome) but uses historical data rather than real-time follow-up.
This design is faster and cheaper than a prospective cohort study but depends on the quality and completeness of existing records.
A researcher identifies 200 patients with a rare autoimmune disease and 200 healthy matched controls, then compares their dietary histories. What type of study is this?
Click to reveal answer
This is a case-control study. The researcher started with the outcome (autoimmune disease vs. no disease) and looked backward at past exposures (diet). Case-control studies are ideal for rare diseases because you can recruit cases directly rather than waiting for them to develop. This study can calculate an odds ratio but not relative risk or incidence.
Why can a case-control study calculate an odds ratio but not a relative risk?
Click to reveal answer
Relative risk requires knowing the incidence of disease in exposed and unexposed groups. In a case-control study, you select participants based on whether they have the disease, so you artificially set the ratio of cases to controls. This means you cannot calculate the actual rate of disease in the exposed population. The odds ratio approximates relative risk when the disease is rare (the rare disease assumption).
Before a new drug reaches your local pharmacy, it goes through a gauntlet of testing that typically takes 10 to 15 years and costs over a billion dollars. Imagine a funnel: thousands of candidate compounds enter at the top, and only a handful survive to the bottom. Each phase of clinical trials is a progressively stricter filter, designed to catch problems before a drug is released to millions of people.
The MCAT does not expect you to know the regulatory fine print of drug approval. What it does expect is that you can identify which phase a passage is describing, explain why each phase exists, and recognize what a randomized controlled trial (RCT) is and why it is considered the gold standard.
The drug development pipeline from discovery to market. Each phase serves as a progressively stricter filter: preclinical testing establishes basic safety, Phase I tests safety in humans, Phase II tests efficacy, Phase III confirms results at scale, and Phase IV monitors long-term effects after approval. Credit: Wikimedia Commons, CC BY-SA 3.0
Before Human Trials: Preclinical Testing
Before any drug is tested in humans, it goes through preclinical testing - laboratory studies (in vitro, meaning in test tubes or cell cultures) and animal studies (in vivo). The goal is to spot obvious toxicity, find safe dosing ranges, and understand how the drug is metabolized.
Only if preclinical results are promising does the drug advance to human trials. An Investigational New Drug (IND) application must be approved before Phase I begins.
Phase I: Safety
Goal: Is the drug safe in humans? What dose is tolerable?
Participants: Small group (20-100), usually healthy volunteers (not patients with the disease).
Duration: Months.
Key question: Does this drug cause unacceptable side effects? What is the maximum tolerated dose?
Phase I is not testing whether the drug works - it is testing whether it is safe enough to keep studying.
Phase II: Efficacy
Goal: Does the drug actually work? What is the optimal dose?
Participants: Moderate group (100-300), patients who have the target condition.
Duration: Months to a couple of years.
Key question: Does this drug produce a measurable therapeutic effect? The study starts to assess efficacy while continuing to monitor safety.
Most drugs fail in Phase II. A drug may be safe but simply not effective enough to justify continued development.
Phase III: Large-Scale Confirmation
Goal: Confirm efficacy in a large, diverse population. Compare to existing treatments or placebo.
Participants: Large group (1,000-5,000+), patients from multiple sites (multicenter).
Duration: One to four years.
Design: This is where the randomized controlled trial (RCT) lives. Phase III trials are typically randomized, double-blind, and placebo-controlled or active-comparator-controlled. They are the gold standard for showing whether a treatment works.
Key question: Does the benefit outweigh the risk in a real-world patient population?
Phase IV: Post-Market Surveillance
Goal: Monitor long-term safety and effectiveness after the drug is on the market.
Participants: General population (thousands to millions of patients using the drug).
Duration: Ongoing, indefinitely.
Key question: Are there rare side effects that only appear in larger populations or with longer use? Should the drug be recalled or have its labeling changed?
Phase IV is critical because some adverse effects are too rare to catch even in Phase III trials. A side effect occurring in 1 in 50,000 patients will not show up in a trial of 3,000 people.
Summary Table
Phase
Goal
Participants
Size
Key Feature
Preclinical
Basic safety
Animals/cells
Varies
No humans
I
Safety, dosing
Healthy volunteers
20-100
First-in-human
II
Efficacy, dosing
Patients with disease
100-300
First test of effectiveness
III
Confirm efficacy
Patients, multicenter
1,000-5,000+
RCT, gold standard
IV
Long-term safety
General population
Thousands+
Post-market surveillance
The Randomized Controlled Trial (RCT)
The RCT deserves special attention because it is the design the MCAT references most often. Its key features:
Randomization: Participants are randomly assigned to treatment or control groups
Control group: Receives placebo or standard-of-care treatment
Blinding: Ideally double-blind (participant and researcher both unaware of assignment)
Prospective: Follows participants forward in time from treatment to outcome
The RCT is the only study design that can establish causation with high confidence, because randomization and controlled conditions remove most confounders.
A new drug showed no serious side effects in 80 healthy volunteers and now moves to testing in 250 patients with the disease. Which phase is this trial entering?
Click to reveal answer
Phase II. The drug has passed Phase I (safety in healthy volunteers) and is now being tested for efficacy in a moderate-sized group of patients who actually have the target condition. Phase II shows whether the drug produces a measurable therapeutic effect.
Why is the randomized controlled trial (RCT) considered the gold standard for establishing causation?
Click to reveal answer
The RCT combines randomization (spreads confounders evenly), a control group (provides a baseline for comparison), blinding (removes placebo effect and observer bias), and prospective design (establishes temporal sequence). Together, these features minimize bias and isolate the effect of the independent variable, making causation the most likely explanation for any observed difference.
Imagine two cooking scenarios. In the first, you follow a recipe perfectly in a professional test kitchen with calibrated equipment - every measurement is exact, every variable is controlled. But you only used one brand of flour that nobody else can buy. Your recipe works flawlessly in that kitchen, but you have no idea if it works anywhere else. In the second scenario, you cook in a real home kitchen with whatever flour is available. Your results are messier, but they are much more representative of what an actual home cook would experience.
The first scenario has high internal validity but low external validity. The second has the opposite. Every study faces this tension, and the MCAT expects you to evaluate both.
Internal Validity
Internal validity asks: “Did the independent variable actually cause the change in the dependent variable?”
A study has high internal validity when you can confidently say the treatment - and nothing else - produced the observed effect. This requires removing confounders, using proper controls, randomizing participants, and blinding where possible.
Threats to internal validity include:
Confounding variables: A third factor explains the result
Selection bias: Groups were not equivalent at baseline
Maturation: Participants change naturally over time (they grow older, more experienced, or heal on their own)
History: External events during the study affect the outcome (e.g., a pandemic starts mid-trial)
Attrition: Participants who drop out differ systematically from those who stay
Testing effects: Taking a pretest changes performance on the posttest
External Validity
External validity asks: “Do the results generalize to other populations, settings, and times?”
A study has high external validity when its findings apply beyond the specific participants, location, and conditions of the study. A drug tested only on 20-year-old male college students may not work the same way in elderly women. A therapy tested in a university lab may not work in a rural clinic.
Threats to external validity include:
Non-representative sample: Only studying one demographic, age group, or geographic region
Artificial setting: Laboratory conditions that do not reflect real-world complexity
Hawthorne effect: Participants behave differently because they know they are being observed (the results may not replicate when observation stops)
Temporal factors: Results from the 1970s may not apply today
The Internal-External Validity Trade-off
There is a natural tension between the two. Tightly controlling an experiment (to boost internal validity) often makes it less realistic (reducing external validity). Studying participants in their natural environment (to boost external validity) introduces confounders (reducing internal validity).
Randomized controlled trials in clinical settings tend to have high internal validity but limited external validity (strict inclusion criteria, controlled conditions). Observational studies in the general population tend to have higher external validity but lower internal validity (no randomization, more confounders).
Construct Validity
Construct validity asks: “Does the measurement tool actually measure the abstract concept it claims to measure?”
This is especially relevant in psychology and social science, where many variables are abstract constructs (intelligence, depression, anxiety, self-esteem).
An IQ test has high construct validity for measuring cognitive ability if it correlates with other accepted measures of intelligence and predicts outcomes that intelligence should predict (academic performance, problem-solving ability).
A “creativity test” that only measures vocabulary has low construct validity for creativity because vocabulary is not the same thing as creativity.
Two subtypes you should recognize:
Convergent validity: The measure correlates with other measures of the same construct (an anxiety questionnaire correlates with physiological stress markers)
Discriminant validity: The measure does not correlate with measures of unrelated constructs (an anxiety questionnaire does not correlate with a math ability test)
Face Validity
Face validity asks: “Does the test look like it measures what it is supposed to measure?”
This is the weakest form of validity - it is just a subjective judgment. A math test that contains math problems has high face validity. A math test that contains only word puzzles has low face validity, even if it turns out to predict math ability surprisingly well.
Face validity matters for participant buy-in (people are more motivated to complete a test that looks relevant), but it does not guarantee that the test actually measures the right thing.
Content and Criterion Validity
Two more types sometimes appear on the MCAT:
Content validity: Does the test cover all aspects of the construct? A biology exam that only tests genetics has low content validity for “biology knowledge” because it ignores ecology, physiology, and cell biology.
Criterion validity: Does the test predict a relevant real-world outcome? A medical school admissions test has high criterion validity if students who score well also perform well in medical school. Criterion validity has two subtypes:
Predictive validity: The test predicts future performance
Concurrent validity: The test correlates with a current gold-standard measure
A drug trial uses strict inclusion criteria (only males aged 25-35, no comorbidities). The drug works. Which type of validity is threatened, and why?
Click to reveal answer
External validity is threatened. The strict inclusion criteria mean the sample is not representative of the broader patient population (women, older adults, patients with other conditions are excluded). While the study may have high internal validity, the results may not generalize to the diverse populations that would actually use the drug.
A researcher uses "number of smiles per hour" to measure happiness. What type of validity is most directly in question?
Click to reveal answer
Construct validity. The question is whether "number of smiles per hour" actually measures the abstract concept of happiness. People may smile for social reasons (politeness, nervousness) without being happy, and genuinely happy people may not smile frequently. A better approach might combine smiling with self-report scales and physiological measures for stronger construct validity.
You step on your bathroom scale three mornings in a row. It reads 160, 160, 160. Great - the scale is consistent. But then you weigh yourself on a calibrated medical scale and it reads 155. Your bathroom scale was reliable (it gave the same answer every time) but not valid (the answer was wrong). Now imagine a different scale that reads 155, 162, 148. This one is neither reliable nor valid. Reliability is about consistency. Validity is about accuracy. You need both for good science, and the MCAT tests the difference.
What Is Reliability?
Reliability is the degree to which a measurement yields consistent, reproducible results. A reliable test gives you the same answer (or very close) every time you use it under the same conditions.
Think of reliability as precision. A reliable thermometer reads 98.6 degrees Fahrenheit every time you check a healthy person’s temperature. It may or may not be accurate (that is validity), but it is consistent.
Test-Retest Reliability
Test-retest reliability measures whether the same test produces the same results when given to the same people at two different time points.
How it works: Give a test at Time 1. Wait a period (days, weeks). Give the same test at Time 2. Correlate the two sets of scores. A high correlation (close to 1.0) means high test-retest reliability.
Example: A depression questionnaire administered today and again two weeks later should produce similar scores (assuming the person’s depression has not actually changed).
Limitation: If the construct being measured naturally fluctuates (mood, anxiety, energy level), low test-retest reliability might reflect genuine change rather than a flawed test.
Inter-Rater Reliability
Inter-rater reliability (also called inter-observer reliability) measures whether different observers produce the same ratings or scores when evaluating the same thing.
How it works: Two or more raters independently rate the same set of participants, essays, or behaviors. You calculate the agreement between raters. High agreement means high inter-rater reliability.
Example: Two psychiatrists independently interview the same 50 patients and assign diagnoses. If they agree on 45 out of 50 diagnoses, inter-rater reliability is high.
Why it matters: If a measurement depends on human judgment (grading essays, rating behavior, reading imaging), inter-rater reliability makes sure the results don’t depend on which specific person does the evaluating.
Common statistics for inter-rater reliability include Cohen’s kappa (for two raters with categorical data) and intraclass correlation coefficient (ICC) (for continuous data or more than two raters).
Internal Consistency
Internal consistency measures whether items within a single test that are supposed to measure the same construct actually produce similar results.
How it works: Look at the correlations among all items on a test. If the items are measuring the same underlying construct, they should correlate with each other.
Cronbach’s α (α) is the most common measure. It ranges from 0 to 1, where values above 0.7 are generally considered acceptable. A depression questionnaire with high α means that people who score high on one question about sadness also tend to score high on other questions about sadness.
Split-half reliability is a simpler version: divide the test items into two halves (odd-numbered vs. even-numbered questions) and correlate the scores from each half. High correlation means high internal consistency.
Reliability Summary Table
Type
Question It Answers
Method
Test-retest
Same results over time?
Give same test twice, correlate scores
Inter-rater
Different observers agree?
Multiple raters score same items, measure agreement
Internal consistency
Items within test agree?
Cronbach’s α or split-half correlation
Reliability vs. Validity: The Critical Relationship
Accuracy (validity) vs. precision (reliability) illustrated with dartboard targets. Reliable but not valid: darts cluster tightly but miss the bullseye. Valid and reliable: darts cluster tightly around the bullseye. A measurement must be reliable (consistent) before it can be valid (accurate). Credit: Wikimedia Commons, CC BY-SA 3.0
This is one of the most tested concepts in research design on the MCAT:
Reliability is NECESSARY but NOT SUFFICIENT for validity.
A test must first be reliable (consistent) before it can be valid (accurate). If a test gives different results every time, it cannot possibly be measuring the right thing - it is not measuring anything consistently. But a test can be perfectly reliable and still not valid - your bathroom scale reads 160 every single time, but you actually weigh 155.
The four possible combinations:
Reliable
Not Reliable
Valid
Ideal: consistent and accurate
Impossible: cannot be accurate if inconsistent
Not Valid
Consistent but wrong (bathroom scale reads 160 every time, but true weight is 155)
Inconsistent and wrong (worst case)
A personality test produces very different scores each time the same person takes it. Is this a problem with reliability, validity, or both?
Click to reveal answer
Both. The test has low test-retest reliability (inconsistent scores across time). Because reliability is necessary for validity, low reliability automatically means the test also lacks validity - it cannot be measuring the right thing if it cannot measure anything consistently.
Two radiologists independently read the same 100 MRI scans. They agree on the diagnosis in 95 out of 100 cases. What type of reliability is this, and is it high or low?
Click to reveal answer
This is inter-rater reliability, and it is high (95% agreement). When different observers produce the same results for the same data, the measurement is reliable across raters. This is especially important for diagnostic imaging, where the interpretation depends on human judgment.
During World War II, the Allied military examined bombers returning from missions and noted where the bullet holes were concentrated - the wings and fuselage. They proposed adding armor to those areas. Statistician Abraham Wald pointed out the flaw: they were only looking at planes that survived. The planes that were hit in the engine or cockpit never made it back. The military was falling victim to survivorship bias - drawing conclusions from an incomplete sample that excluded the most important cases.
Bias is any systematic error that skews results in a particular direction. Unlike random error (which averages out with large samples), bias pulls every measurement the same wrong way. The MCAT expects you to identify specific types of bias, explain how they distort a study’s conclusions, and suggest how to fix them.
Selection Bias
Selection bias occurs when the sample is not representative of the population, usually because of how participants were recruited or assigned.
Example: A study on exercise habits recruits participants from a gym. People who go to gyms already exercise more than the general population, so the results overestimate how much the average person exercises.
How to reduce it: Use random selection from the target population. Use random assignment to spread baseline differences across groups.
Recall Bias
Recall bias happens when participants inaccurately remember past events, and the inaccuracy is systematic. People with a disease may think harder about past exposures (“I must have been exposed to something”) while healthy controls may not.
Example: In a case-control study of childhood leukemia, parents of children with leukemia may recall environmental exposures (power lines, chemicals) more readily than parents of healthy children, inflating the apparent association.
How to reduce it: Use prospective designs (collect data before the outcome happens). Use medical records or other objective data instead of relying on memory.
Attrition Bias
Attrition bias (dropout bias) happens when participants who leave a study differ systematically from those who stay. If sicker patients drop out of a drug trial because of side effects, the remaining participants look healthier, making the drug appear more effective than it really is.
Example: A weight loss study starts with 200 participants. By the end, 50 have dropped out - mostly those who did not lose weight. The final results show impressive average weight loss, but only because the failures left the data set.
How to reduce it: Track reasons for dropout. Use intention-to-treat analysis, which includes all participants in the final analysis whether or not they completed the study.
Hawthorne Effect
The Hawthorne effect happens when participants change their behavior simply because they know they are being observed, regardless of the experimental treatment.
Example: Workers in a factory increase their productivity when researchers observe them - not because of any change in working conditions, but because being watched motivates them to perform better.
How to reduce it: Use blinding (so participants do not know they are being studied). Use unobtrusive measures. Include a control group that is also observed.
Observer (Experimenter) Bias
Observer bias happens when the researcher’s expectations influence how they collect, interpret, or record data.
Example: A researcher who believes a drug works may unconsciously rate treated patients as more improved, measure outcomes more favorably, or probe for positive responses.
How to reduce it: Double-blind design (the researcher does not know which participants are in the treatment group). Use objective, standardized measurement tools.
Response Bias
Response bias is a broad category where participants do not answer truthfully. It includes:
Social desirability bias: Participants give answers they think are socially acceptable (“I exercise daily” when they do not)
Acquiescence bias: Participants tend to agree with statements regardless of content (yes-saying)
Demand characteristics: Participants figure out the study’s hypothesis and adjust their behavior to confirm (or contradict) it
How to reduce it: Use anonymous surveys. Include reverse-coded items. Avoid leading questions.
Confirmation Bias
The cognitive bias codex organizes the many known cognitive biases into categories. For the MCAT, the most relevant are confirmation bias, recall bias, selection bias, and the Hawthorne effect. Awareness of these biases helps researchers design better studies and helps readers evaluate published findings critically. Credit: Wikimedia Commons, CC BY-SA 4.0
Confirmation bias is the tendency to seek, interpret, and remember information that confirms pre-existing beliefs while ignoring contradictory evidence.
In research: A scientist who believes their hypothesis is correct may selectively report supportive data, dismiss contradictory findings as “outliers,” or design experiments that can only confirm (never disconfirm) the hypothesis.
How to reduce it: Preregister hypotheses and analysis plans. Use peer review. Actively seek disconfirming evidence.
Survivorship Bias
Survivorship bias happens when you draw conclusions from only the “survivors” - the cases that made it through a selection process - while ignoring those that did not.
The bomber example from the introduction is the classic case. In medicine: if you study long-term outcomes of a disease by surveying patients in a clinic, you are only seeing patients who survived long enough to be in the clinic. The sickest patients already died and are invisible to your study.
In everyday life: “College dropouts like Bill Gates became billionaires, so dropping out must be fine” ignores the millions of dropouts who did not become billionaires.
Bias Summary Table
Bias
Source
Direction of Distortion
Selection
Non-representative sample
Results do not apply to target population
Recall
Inaccurate memory of past events
Inflates or deflates apparent associations
Attrition
Differential dropout
Remaining sample is biased toward those who respond well
Hawthorne
Awareness of being observed
Behavior improves artificially
Observer
Researcher expectations
Outcomes measured or interpreted favorably
Response
Participant dishonesty or acquiescence
Self-report data skewed
Confirmation
Pre-existing beliefs
Evidence selectively interpreted
Survivorship
Missing the “non-survivors”
Overestimates success or underestimates risk
In a drug trial, 30% of the treatment group drops out due to side effects. The remaining participants show significant improvement. What bias is present?
Click to reveal answer
Attrition bias. The participants who dropped out likely experienced the worst outcomes (side effects, lack of improvement). The remaining participants are a biased sample of those who tolerated the drug well. The drug appears more effective than it truly is because the failures have been removed from the data. An intention-to-treat analysis would address this by including all participants in the final results.
Parents of children with autism are asked to recall their child's early vaccination history. Parents of healthy children are asked the same questions. Which bias is most likely to affect the results?
Click to reveal answer
Recall bias. Parents of children with autism may have spent more time thinking about possible causes and may recall vaccinations (or perceived reactions) more vividly and in more detail than parents of healthy children. This differential recall can artificially inflate the association between vaccination and autism in a case-control study.
In the 1930s, researchers from the U.S. Public Health Service began studying the natural progression of syphilis in a group of Black men in Tuskegee, Alabama. The men were told they were receiving free treatment for “bad blood.” In reality, they received no treatment at all - even after penicillin became the standard cure in the 1940s. The study continued for 40 years. Men died, went blind, and passed the infection to partners and children. When the study was exposed in 1972, the public outrage led directly to modern research ethics regulations.
Research ethics exist because the history of science includes genuine atrocities. The MCAT tests your knowledge of the ethical principles and oversight mechanisms that now protect human subjects - and occasionally presents passages based on historical cases like Tuskegee.
Informed Consent
Informed consent means that participants must be given all relevant information about a study and must voluntarily agree to participate. It is not a form you sign once - it is an ongoing process.
The core elements of informed consent:
Purpose of the study and procedures involved
Risks and benefits of participation
Voluntary participation - the right to withdraw at any time without penalty
Confidentiality - how data will be protected
Contact information for questions or concerns
Consent must be given freely, without coercion. A prisoner who is offered early release for participating is not giving truly voluntary consent. A patient who is told by their doctor to “definitely enroll” is under undue influence.
Institutional Review Board (IRB)
The Institutional Review Board (IRB) is an independent committee that reviews and approves all research involving human subjects before the study begins. Every university, hospital, and research institution has one.
The IRB evaluates whether:
Risks to participants are minimized and reasonable relative to potential benefits
Informed consent procedures are adequate
Participant selection is fair (not targeting vulnerable populations unfairly)
Data is kept confidential
Extra safeguards exist for vulnerable populations
No human subjects research can proceed without IRB approval. Even changes to an already-approved study must be reviewed.
The Belmont Report and Four Principles
The Belmont Report (1979), written in response to the Tuskegee scandal, established three core principles for research ethics. The MCAT often tests these alongside a fourth principle from biomedical ethics:
1. Respect for Persons (Autonomy): Individuals are autonomous agents and should make their own decisions about participation. Those with diminished autonomy (children, cognitively impaired individuals) deserve extra protections.
2. Beneficence: Researchers must maximize potential benefits and minimize potential harms. This is a positive duty - you must actively do good, not just avoid harm.
3. Non-maleficence: “First, do no harm.” Researchers must not expose participants to unnecessary risk. This is the negative duty - do not cause harm.
4. Justice: The benefits and burdens of research must be spread fairly. Vulnerable populations should not bear a disproportionate share of research risk while others receive the benefits.
Vulnerable Populations
Certain groups need extra ethical protections because their ability to give truly voluntary, informed consent is compromised:
Children: Cannot consent; require parental permission (consent) plus the child’s assent
Prisoners: Institutional setting creates coercive pressure; any incentives may be unduly influential
Pregnant women: Research may affect the fetus, who cannot consent
Cognitively impaired individuals: May not fully understand the risks and procedures
Economically disadvantaged people: Financial incentives may constitute undue inducement
Deception in Research
Sometimes telling participants the true purpose of a study would invalidate the results. If participants in a conformity study know they are being tested on conformity, they will behave differently.
Deception is permissible only when:
The study cannot be done without it
The research has significant scientific or educational value
No reasonable alternative exists
Participants are not exposed to significant risk
Participants are debriefed afterward - told the true purpose, why deception was necessary, and given the chance to withdraw their data
Debriefing is mandatory after any study involving deception. The researcher must explain what was actually being studied, why the deception was necessary, and answer any questions.
Confidentiality and Privacy
Researchers must protect participants’ personal information. Data should be anonymized or de-identified whenever possible. If data cannot be anonymized, access should be restricted and data should be stored securely.
Anonymity means the researcher cannot link data to individual participants (no names, no identifying codes). Confidentiality means the researcher can link data to individuals but promises not to disclose it.
A researcher conducts a study where participants are told they are testing memory, but the study actually measures racial bias. Which ethical requirement must be met after the study?
Click to reveal answer
Debriefing is required. After a study involving deception, the researcher must explain the true purpose of the study, explain why deception was necessary, and give participants the chance to withdraw their data. Deception is only permissible when the study cannot be conducted without it and when participants are not exposed to significant risk.
A clinical trial offers prisoners reduced sentences for participation. Which ethical principle is most directly violated?
Click to reveal answer
Respect for persons (autonomy) is violated. Prisoners are a vulnerable population, and offering sentence reduction creates coercive pressure that undermines truly voluntary consent. The justice principle is also at risk, because a disproportionate burden of research risk falls on an already disadvantaged group.
A news article claims that a single study “proves” red wine prevents cancer. A supplement company cites “published research” showing their product boosts memory by 300%. A viral social media post links a journal paper as evidence that a common food is toxic. In each case, the claim sounds scientific. But the MCAT trains you to look behind the curtain: Was the study peer-reviewed? Has it been replicated? Is the effect statistically significant and practically meaningful? Does the journal even have editorial standards?
Evaluating scientific claims is the capstone skill of this chapter. Everything you have learned about study design, variables, controls, validity, and bias converges here. The MCAT presents passages from real or realistic research studies, and your job is to assess the quality of the evidence and the strength of the conclusions.
Peer Review
The hierarchy of research evidence. Systematic reviews and meta-analyses sit at the top, providing the strongest evidence. Randomized controlled trials are below, followed by cohort and case-control studies. Expert opinion and case reports provide the weakest evidence. Stronger designs minimize bias and confounding. Credit: Wikimedia Commons, CC BY-SA 3.0
Peer review is the process by which other experts in the field evaluate a study’s methods, analysis, and conclusions before it is published in a journal.
How it works: A researcher submits a manuscript to a journal. The editor sends it to two or three independent experts (peers) who were not involved in the study. The reviewers critique the methodology, point out flaws, and recommend whether to accept, revise, or reject the paper. The process is usually anonymous (the reviewers do not know the authors, or vice versa, depending on the journal).
Why it matters: Peer review is the primary quality filter in science. It is not perfect - reviewers can miss errors, hold biases, or be slow - but it is far better than no review at all. A peer-reviewed study has at least been scrutinized by experts, which gives it more credibility than a non-reviewed preprint, blog post, or press release.
Replication and Reproducibility
Replication means running the same study again (ideally by a different research team) to see if the results hold up. A finding that cannot be replicated is far less trustworthy than one confirmed by multiple independent groups.
Reproducibility is a related but slightly different concept: can another researcher, given the same data and methods, arrive at the same conclusions? Reproducibility is about the analysis; replication is about the entire experiment.
The replication crisis refers to the widespread finding that many published results - particularly in psychology and biomedical science - fail to replicate when other teams attempt to repeat them. This has raised awareness about the importance of larger sample sizes, preregistration, and transparent reporting.
Statistical Significance vs. Practical Significance
Statistical significance means the result is unlikely to have happened by chance alone. The conventional threshold is p < 0.05, meaning there is less than a 5% probability of observing this result if the null hypothesis were true.
Practical significance (also called clinical significance) means the result is large enough to matter in the real world.
These two can diverge:
Statistically significant but not practically significant: A study with 100,000 participants finds that a new drug lowers blood pressure by 0.5 mmHg (p < 0.001). The result is real (not due to chance) but clinically meaningless - no doctor would prescribe a drug for half a millimeter of mercury.
Practically significant but not statistically significant: A study with 15 participants finds that a drug lowers blood pressure by 20 mmHg (p = 0.08). The effect size is huge, but the small sample size means the study lacks the statistical power to reach p < 0.05.
Effect Size
Effect size measures how big the difference between groups is, independent of sample size. Common measures include Cohen’s d (for comparing two means) and the correlation coefficient r.
Effect size matters because p-values are heavily influenced by sample size. A trivially small effect can reach p < 0.05 with a large enough sample. Effect size tells you how big the difference is, not just whether it exists.
Limitations Sections
Every well-written study includes a limitations section that honestly acknowledges weaknesses. Common limitations include:
Reliance on self-report data (subject to response bias)
Short follow-up period (may miss long-term effects)
Lack of blinding or randomization (reduces internal validity)
On the MCAT, passages sometimes include a study’s limitations. Questions may ask you to identify additional limitations the authors did not mention, or to evaluate whether a stated limitation actually affects the conclusions.
Publication Bias
Publication bias (also called the file-drawer problem) happens because studies with positive, statistically significant results are much more likely to be published than studies with null or negative results. The studies that “did not work” sit in file drawers, unpublished.
This creates a distorted picture of reality. If 20 teams independently test a drug and 19 find no effect but 1 finds a positive result (by chance), only the positive result gets published. Anyone reading the literature sees 100% positive evidence, when in reality the drug probably does not work.
Strategies to combat publication bias:
Preregistration: Researchers publicly register their hypothesis and methods before collecting data, making it harder to hide null results
Registered reports: Journals agree to publish the study based on its design, regardless of results
Meta-analyses: Statistical methods that combine results from multiple studies, sometimes including unpublished data
How to Read an Abstract on the MCAT
MCAT science passages often resemble research paper abstracts. Here is how to read them efficiently:
Identify the research question. What are they trying to find out?
Classify the study type. Experimental? Observational? Which subtype?
Identify the variables. What is the IV? DV? Were confounders controlled?
Evaluate the design. Controls? Blinding? Sample size? Randomization?
Assess the conclusion. Does it follow from the data? Does it overextend?
You do not need to understand every detail of the methods. Focus on the elements this chapter has taught you: variables, controls, design type, bias risks, and whether the conclusion is supported.
A study of 50,000 participants finds that a supplement improves memory by 0.1% (p = 0.002). Should a doctor recommend this supplement?
Click to reveal answer
Probably not. The result is statistically significant (p = 0.002, well below 0.05), so the improvement is likely real and not due to chance. However, it is not practically significant - a 0.1% improvement in memory is too small to be meaningful in clinical practice. The large sample size (50,000) gave the study enough power to detect a trivially small effect.
Why does publication bias distort the scientific literature?
Click to reveal answer
Publication bias distorts the literature because studies with positive results are more likely to be published than studies with null or negative results. This means the published evidence overrepresents positive findings, creating a misleading impression that a treatment works or an association exists. Someone reviewing the literature sees only the "successes" while the null results remain unpublished in file drawers.