Memory, Attention, and Cognition

Chapter 4: Memory, Attention, and Cognition

Full chapter view · 7 sections · ~41 min read Switch to section-by-section view →
4.1

Attention and Selective Attention

Your eyes take in about 10 million bits of information per second. Your conscious attention processes about 40. The difference is what this section is about: attention is the ferocious filter that decides which tiny fraction of the sensory flood actually reaches awareness.

Types of Attention

  • Selective attention. Focusing on one stimulus while ignoring others. Following one voice at a party.
  • Divided attention. Trying to handle multiple attention-demanding tasks simultaneously. Driving while texting.
  • Sustained (directed) attention. Maintaining focus on a single task over time. Reading this page.
  • Alternating attention. Switching between tasks, one at a time.

Attention is a limited resource. Doing two demanding things at once means either rapidly switching between them (not actually doing both simultaneously) or doing both poorly. The popular concept of “multitasking” is mostly task-switching, and it comes with a cost: each switch takes time and effort.

Two rows of color words. In the top row each word is printed in its matching ink color (Red in red, Green in green, Purple in purple, etc.). In the bottom row each word is printed in a non-matching color (Purple in green ink, Red in brown ink, Brown in blue ink, etc.) — naming the ink color is much slower for the bottom row
The Stroop task. Naming the ink color of the top row is fast and automatic; naming the ink color of the bottom row is slow because the word's meaning interferes — a classic demonstration of how automatic reading consumes attention you wanted for color naming. Credit: HouseBlaster & Belbury via Wikimedia Commons (Public Domain).

Inattentional Blindness and Change Blindness

Two classic failures of attention:

  • Inattentional blindness - failing to see an unexpected stimulus that is right in your visual field because your attention is focused elsewhere. The famous “invisible gorilla” experiment: viewers counting basketball passes fail to notice a person in a gorilla suit walking through the scene.
  • Change blindness - failing to notice changes between a previous and current state of a scene. Classic demonstration: a person asking directions from a stranger is briefly blocked, swapped for a different person, and the stranger doesn’t notice.

Both phenomena prove that perception is not a passive video recording - it is a selective reconstruction gated by attention.

The Cocktail Party Effect

Cocktail party effect: the ability to focus on one voice in a noisy crowd while still detecting salient signals (like your own name) in unattended streams. This observation drove decades of theories about how selective attention works.

Theories of Selective Attention

The classic dichotic listening experiment plays one message into each ear and asks the subject to repeat (shadow) the message in one ear. What happens to the unattended message?

Broadbent’s Early Selection (1958)

All sensory information enters a sensory buffer. A selective filter operates immediately, based on physical features (pitch, location, voice). Only one message gets passed on to perceptual analysis and meaning extraction. The unattended message is blocked.

Pathway: sensory register → selective filter → perceptual processing → consciousness.

Problem: If the unattended message is truly blocked, how do you hear your name in the other ear (cocktail party effect)?

Deutsch and Deutsch’s Late Selection

Everything gets processed for meaning. The filter comes after perception; it decides what reaches consciousness.

Pathway: sensory register → perceptual processing → selective filter → consciousness.

Problem: This is computationally expensive - processing everything for meaning seems wasteful.

Treisman’s Attenuation Theory

A middle ground. Instead of a hard filter, there is an attenuator that weakens (not eliminates) the unattended stream. All streams still reach perception, but the attended one is loud and the others are whispered. Important words (your name, the word “fire”) can still break through if loud enough relative to their attenuated volume.

Pathway: sensory register → attenuator → perceptual processing → consciousness.

Treisman’s theory handles the cocktail party effect elegantly and is the most widely cited of the three on the MCAT.

Spotlight and Resource Models of Attention

  • Spotlight model. Attention is like a movable spotlight that illuminates one small area of the perceptual field at a time. You can redirect it, but only one area is “lit” at a time.
  • Resource model. Attention is a limited pool of cognitive resources. Tasks consume resources; when demand exceeds supply, performance drops. Explains why doing two demanding things at once is worse than doing either alone, even when they use different sensory modalities.

Multitasking: Task Similarity, Difficulty, Practice

Three factors determine how well you can do two things at once:

  • Task similarity. Two tasks that use the same cognitive system (both verbal, both visual) interfere more than tasks using different systems. You can listen to classical music while writing a paper; you cannot listen to someone’s speech while writing.
  • Task difficulty. Harder tasks demand more attention. Driving in heavy traffic leaves no room for a phone call; driving on an empty highway leaves plenty.
  • Practice. Skilled tasks become automatic and demand little attention. Adult drivers can hold a conversation while driving, something impossible for new drivers. The transition from controlled to automatic processing is the core benefit of practice.
What is inattentional blindness?
Click to reveal answer
Failing to see an unexpected stimulus directly in your visual field because your attention is focused elsewhere. Famous demo: the invisible gorilla in the basketball-counting video.
Which of the three classic selective attention theories best accounts for the cocktail party effect?
Click to reveal answer
Treisman's attenuation theory. The unattended stream is weakened, not blocked, so high-salience words (your name, "fire") can still break through. Broadbent's strict filter cannot easily explain the effect.
Why can skilled drivers chat while driving while new drivers cannot?
Click to reveal answer
Practice converts driving from controlled processing (attention-intensive) to automatic processing (near-effortless). Automatic tasks free up attentional resources for secondary tasks like conversation.
4.2

The Information-Processing Model of Memory

The information-processing model treats the brain like a computer: INPUT → PROCESS → OUTPUT. Information flows through three stages of memory: sensory, working (short-term), and long-term. Each has a different capacity, duration, and purpose. This is the scaffolding of everything in the memory section.

Flowchart of the Atkinson-Shiffrin multi-store model: environmental input enters sensory memory, is filtered by attention into short-term memory, where rehearsal can either maintain it in STM or transfer it via elaborative rehearsal into long-term memory; retrieval flows the other direction. Arrows label decay, displacement, interference, and retrieval failure as the failure modes for each store
The Atkinson–Shiffrin multi-store model — the architectural backbone of every memory question. Sensory memory holds raw input briefly; attention filters a fraction into short-term memory; elaborative rehearsal moves it into long-term memory. Credit: Dkahng via Wikimedia Commons (CC BY-SA 4.0).

Sensory Memory: The Brief Snapshot

Sensory memory (or sensory register) is the very first stage. It holds an extremely brief, high-capacity snapshot of raw sensory input. Two sub-types, one per modality:

  • Iconic memory - visual sensory memory. Lasts about 0.5 seconds.
  • Echoic memory - auditory sensory memory. Lasts about 3–4 seconds.

Why so brief? Sensory memory is a holding pen that lets the brain decide which information to attend to and promote to working memory. Everything else decays.

Sperling’s experiment (1960) is the classic demonstration. Subjects see 12 letters arranged in a 3×4 grid for 50 ms. In the whole-report condition, they try to name as many letters as possible - typically 3–5 (~35%). In the partial-report condition, an auditory cue after the display tells them which row to report; performance jumps to about 75% per row, implying the whole 12-letter display was briefly available in memory, but faded before it could be reported.

Working Memory (Short-Term Memory)

Working memory holds the information you are currently consciously thinking about. Classic capacity: 7 ± 2 items (Miller, 1956). This is why phone numbers are 7 digits. Items decay within about 20–30 seconds unless rehearsed.

Baddeley and Hitch (1974) proposed a more detailed model with four components:

  • Phonological loop - holds verbal/acoustic information. Capacity around 2 seconds’ worth of sound. Repeating a phone number to yourself uses the phonological loop.
  • Visuospatial sketchpad - holds visual and spatial information. Mental images, mental rotation.
  • Episodic buffer - integrates information from the loop, sketchpad, and long-term memory into a coherent episode. Added to the model in 2000.
  • Central executive - directs attention among the other components. Decides what to focus on, coordinates rehearsal, manages task switching. Think of it as the conductor of the memory orchestra.
Diagram of Baddeley and Hitch's working memory model. A central executive sits at the top with bidirectional arrows to three subsystems below: the phonological loop (handling language), the visuospatial sketchpad (handling visual semantics), and the episodic buffer (handling short-term episodic memory)
Baddeley's working memory model. The central executive directs three "slave" subsystems: the phonological loop holds verbal/acoustic information, the visuospatial sketchpad holds visual and spatial information, and the episodic buffer integrates them into a coherent episode. Credit: Mirek2 via Wikimedia Commons (CC0).

Measuring Working Memory: Span Tasks

Working memory span tasks quantify how much information a person can actively hold and manipulate. They are heavily used in research and clinical assessment. Three variants you should know:

  • Digit span (simple span). The examiner reads out a string of digits; the participant repeats them back. Measures the storage side of working memory only. Normal adult span: about 7 ± 2.
  • Backward digit span. Same task, but the participant must repeat the digits in reverse order. Requires simultaneously holding the digits and reorganizing them. Measures both storage and manipulation - central-executive work. Typical span drops to 5–6.
  • Operational span (OSPAN). A dual-task paradigm. The participant solves a short math problem, then reads a word, then solves another math problem, then reads another word, and so on. At the end of the set they recall the words in order. The math problems force continuous central-executive engagement; the word recall measures residual storage capacity. OSPAN scores predict performance on complex tasks like reading comprehension, multitasking, and fluid intelligence better than simple span does - because real-life cognition almost always involves holding information while doing something else.

Long-Term Memory: Where Things Live Long-Term

Long-term memory (LTM) has effectively unlimited capacity and can last a lifetime. Two broad categories:

Explicit (Declarative) Memory

Memories you can consciously recall and describe. Two sub-types:

  • Episodic memory - autobiographical events. “The time I broke my arm at summer camp.”
  • Semantic memory - general knowledge and facts. “The capital of France is Paris.” “H2O is water.”

Explicit memory is heavily dependent on the hippocampus for encoding new memories. Damage to the hippocampus (famously, patient H.M.) produces severe anterograde amnesia for explicit material but preserves implicit memory.

Implicit (Non-Declarative) Memory

Memories that influence behavior without conscious recall. Several sub-types:

  • Procedural memory - how to do things. Riding a bike, typing, playing a learned piano piece. Stored in the basal ganglia and cerebellum.
  • Priming - prior exposure to a stimulus influences later responses, without awareness.
  • Classical conditioning associations - we saw these in Chapter 3. Unconscious stimulus pairings.

A famous example of the dissociation: patient H.M. could learn a new motor skill (procedural) across days even though he had no conscious memory of practicing. His implicit learning system worked; his explicit was shattered.

Memory TypeConscious?ExampleBrain Region
EpisodicYesMy last birthdayHippocampus, cortex
SemanticYesParis is in FranceTemporal cortex
ProceduralNoRiding a bikeBasal ganglia, cerebellum
PrimingNoRecent word influences later choiceCortex

Long-Term Potentiation: How Memories Stick

Long-term potentiation (LTP) is the cellular mechanism behind memory formation. When a presynaptic neuron repeatedly stimulates a postsynaptic neuron, the synapse strengthens: the same stimulation later produces a bigger postsynaptic response.

Mechanism, briefly:

  • Glutamate released from presynaptic neuron activates AMPA and NMDA receptors on the postsynaptic neuron.
  • NMDA receptors, normally blocked by Mg²⁺, open only when the postsynaptic cell is already depolarized. So they detect coincidence: “pre-synaptic fired AND post-synaptic is excited.”
  • Ca²⁺ flows through NMDA channels, triggering signaling cascades that insert more AMPA receptors and strengthen the synapse.

The motto: “Cells that fire together, wire together” (Hebb’s rule). LTP is the leading candidate neural substrate for learning and memory. It exemplifies synaptic plasticity - the ability of synapses to change their strength with experience.

Flowchart of late long-term potentiation: NMDA, AMPA, and mGluR receptors at the top feed into kinases (PI-3K, PKA, PKC, CaMKII), which converge on ERK; ERK then drives signaling, cytoskeletal, and nuclear protein cascades that produce gene transcription, protein synthesis, and morphological changes — yielding LTP expression at the bottom
The signaling cascade behind late-phase LTP. NMDA receptor activation triggers Ca²⁺ entry, which engages a network of kinases (CaMKII, PKA, PKC) and ultimately drives gene transcription and synaptic remodeling — converting a transient excitation into a structurally stronger synapse. Credit: User:Diberri via Wikimedia Commons (CC BY-SA 3.0).
Name the four components of Baddeley's working memory model.
Click to reveal answer
Phonological loop (verbal), visuospatial sketchpad (visual/spatial), episodic buffer (integrates inputs), and central executive (directs attention among them).
Distinguish episodic, semantic, and procedural memory.
Click to reveal answer
Episodic = autobiographical events you can consciously recall. Semantic = general facts and knowledge. Both are explicit (conscious). Procedural = how to do things (bike riding), implicit (unconscious), stored in basal ganglia.
What does the operational span (OSPAN) task measure, and why is it more predictive than simple digit span?
Click to reveal answer
OSPAN is a dual-task paradigm: solve math problems while remembering words for later recall. It measures working memory capacity under concurrent load, engaging the central executive. Predicts fluid intelligence and reading comprehension better than simple digit span, which measures only passive storage.
What is long-term potentiation and why does it matter?
Click to reveal answer
LTP is the strengthening of a synapse after repeated stimulation: the same presynaptic input produces a larger postsynaptic response. It is the leading cellular model for how memories are stored, driven largely by NMDA receptor activation and AMPA receptor insertion.
4.3

Encoding, Retrieval, and Forgetting

Memory is not a recording device. Getting information in (encoding) and getting it back out (retrieval) are two separate, fallible processes, and forgetting happens at many stages between them. Understanding where failure happens is half the MCAT battle.

Encoding Strategies

Encoding moves information from working memory into long-term storage. Some strategies work better than others:

  • Rote rehearsal. Repeating the same material over and over. The weakest strategy. Sufficient to keep info in working memory, poor for LTM.
  • Chunking. Grouping individual items into meaningful units. The number 1776202520 is ten items; “1776, 2025, 20” is three chunks. Expands working memory capacity effectively.
  • Self-referencing. Tying new material to your own experiences. Relating a historical event to something that happened in your life. Engages rich associative networks.
  • Elaborative rehearsal. Connecting new information to existing knowledge. Explaining it in your own words. Far more effective than rote.
  • Mnemonics.
    • Method of loci - visualize walking through a familiar place and “dropping” items at specific locations. Used by memory champions.
    • Peg word system - link items to a pre-memorized rhyming list (1=bun, 2=shoe, 3=tree…).
    • Acronyms - HOMES for the Great Lakes.
  • Spacing effect. Distributing study over time beats cramming the same total amount into one session. Spaced study is one of the most robust findings in all of cognitive psychology.
  • Encoding specificity. Memory is best when retrieval conditions match encoding conditions. Study in a quiet room, test in a quiet room.

Retrieval Cues

Retrieval depends on cues that were present at encoding.

  • Context-dependent memory. Being in the same physical environment as encoding helps. Scuba divers who learned material underwater remembered it better underwater than on land.
  • State-dependent memory. Being in the same internal state (mood, intoxication, fatigue) at retrieval as at encoding helps. Learn something drunk, easier to recall drunk. Learn something while sad, easier to recall while sad (a nasty feature of depression).
  • Priming. Prior exposure to related concepts eases retrieval of linked ones. Reading about apples speeds up recognizing the word “fruit.”

Types of Retrieval

  • Free recall - produce the material with no cues. Hardest.
  • Cued recall - produce the material with a partial cue (“pl____” for “planet”). Easier.
  • Recognition - pick the correct item from a set (“was it fork or spoon?”). Easiest.

Serial Position Effect

Given a list to remember, people show the serial position curve:

  • Primacy effect - superior recall for items at the beginning of the list. Because you had time to rehearse them into LTM.
  • Recency effect - superior recall for items at the end of the list. Because they are still in working memory.
  • Middle items get the worst of both.

Delay or interference between presentation and recall wipes out the recency effect but not primacy.

A line graph showing percentage of words recalled on the y-axis against position in the studied sequence on the x-axis. The curve is high at the start (Primacy), drops to a minimum in the middle (intermediate, shaded gray), and rises sharply again at the end (Recency), forming a U-shape
The serial position curve. Items at the start of a list are remembered well (primacy effect — they had time to enter long-term memory) and items at the end are remembered well (recency effect — they are still in working memory). Middle items get the worst of both. Credit: Obli via Wikimedia Commons (CC BY-SA 3.0).

Why We Forget

Decay

Memories fade with time and disuse. Ebbinghaus’s forgetting curve: most forgetting happens in the first day; the curve levels off afterward. Relearning is faster than first learning, demonstrating a savings effect - some residue of the memory persists even when you cannot explicitly retrieve it.

A line graph showing memory retention on the y-axis vs days on the x-axis. The red curve drops sharply in the first day and continues to decay slowly. Stacked above are several green curves at progressively higher starting points (days 1, 2, 3) representing relearning sessions, each decaying more slowly than the original
Ebbinghaus's forgetting curve. After a single study session (red), most forgetting happens in the first 24 hours. Each subsequent review (green) starts higher and decays more slowly — the empirical case for spaced repetition. Credit: Icez via Wikimedia Commons (Public Domain).

Interference

Other memories compete with the one you want.

  • Retroactive interference. New learning impairs old material. Learning a new phone number makes the old one harder to recall.
  • Proactive interference. Old learning impairs new material. Habitual old passwords interfere with learning the new password.

Memory Is Reconstructed, Not Replayed

Every time you retrieve a memory, you rebuild it. Reconstruction is imperfect and sometimes flat wrong.

  • False memories. Subjects shown a video of a car at a yield sign, then given a misleading verbal description mentioning a stop sign, frequently remember a stop sign.
  • Misleading-questions effect (Loftus). Asking subjects “how fast were the cars going when they SMASHED into each other?” produces higher speed estimates and more reports of broken glass than asking “when they HIT each other.” The verb loaded the question and reshaped the memory.
  • Source monitoring errors. Forgetting where a memory came from. Recalling a detail from a movie but thinking it was from real life. Source amnesia is the extreme form.

Flashbulb Memories

Emotional events create vivid, detailed memories that feel exact - where you were when you heard about 911\frac{9}{11}, for example. Despite the vividness, flashbulb memories are just as susceptible to reconstruction errors as any other memory. They feel accurate because the emotional tag makes them easy to retrieve and easy to trust, not because the details are protected from distortion.

Amnesia

  • Retrograde amnesia. Inability to recall memories formed before the brain injury or onset. “Retro = before.”
  • Anterograde amnesia. Inability to form new memories after the injury. “Antero = after.”

Both can coexist, and both typically reflect damage to the medial temporal lobe (hippocampus).

Alzheimer’s Disease

Alzheimer’s is the most common form of dementia (the umbrella term for significant cognitive decline interfering with daily life). Neurons die; the cerebral cortex shrinks.

  • Hallmark pathology: amyloid plaques (beta-amyloid aggregates) and neurofibrillary tangles (tau protein).
  • Earliest symptom: difficulty forming new memories (anterograde amnesia) - short-term memory fails first.
  • Progression: episodic, semantic, and eventually procedural memory all decline; language deteriorates; patients may fail to recognize family.
  • Cause unknown. Treatment is symptomatic, not curative. Terminal.
High-magnification histopathology slide of brain tissue showing a roughly circular dense brown amyloid plaque deposit surrounded by stained pink and purple neuropil — the classic extracellular β-amyloid aggregate seen in Alzheimer's disease
An amyloid plaque (the dense central deposit) in cortex from a patient with Alzheimer's disease. Plaques accumulate extracellularly years before clinical symptoms appear and remain a diagnostic hallmark on autopsy. Credit: Mikael Häggström, M.D. via Wikimedia Commons (CC0).

Korsakoff Syndrome

Caused by thiamine (vitamin B1) deficiency, most often from chronic alcoholism. Preceded by Wernicke’s encephalopathy (confusion, abnormal eye movements, ataxia), which if untreated can progress to full Korsakoff syndrome.

  • Severe anterograde and retrograde amnesia.
  • Confabulation - patients invent plausible but false memories to fill the gaps, often without realizing it.
  • Unlike Alzheimer’s, Korsakoff is NOT progressive if treated with thiamine and abstinence. Often stable or improving with proper care.
A test taken in the same classroom where you studied performs better than a test in a new classroom. What principle is at work?
Click to reveal answer
Context-dependent memory (a form of encoding specificity). Matching environmental cues at encoding and retrieval improves performance.
Differentiate retroactive and proactive interference.
Click to reveal answer
Retroactive: new info disrupts retrieval of OLD info (new address makes you forget the old address). Proactive: old info disrupts learning of NEW info (old password keeps popping up when you try to use the new one).
What is the difference between retrograde and anterograde amnesia?
Click to reveal answer
Retrograde: loss of memories formed BEFORE the injury (retro = back). Anterograde: inability to form NEW memories AFTER the injury (antero = forward). Both can occur together, especially with hippocampal damage.
What vitamin deficiency causes Korsakoff syndrome, and what is its hallmark clinical feature?
Click to reveal answer
Thiamine (B1) deficiency, typically due to chronic alcoholism. Hallmark feature is severe memory loss with confabulation (invented filler "memories"). Preceded by Wernicke's encephalopathy.
4.4

Piaget, Schemas, and Problem Solving

Jean Piaget was a Swiss biologist who, early in the 20th century, argued something radical for his time: children are not miniature adults. Their minds do not merely accumulate more knowledge; they progress through qualitatively different stages of thinking. Every MCAT psych passage about cognitive development in children draws from Piaget’s framework.

Black-and-white portrait of an elderly Jean Piaget wearing wire-rimmed glasses and a beret, photographed at the University of Michigan in the 1960s
Jean Piaget (1896–1980) at the University of Michigan, c. 1968. His four-stage theory of cognitive development still anchors every MCAT developmental-psych question. Credit: 1968 Michiganensian (University of Michigan), Public Domain via Wikimedia Commons.

Piaget’s Four Stages of Cognitive Development

Memorize the ages and the landmark achievement of each stage.

Stage 1: Sensorimotor (0–2 years)

Infants explore the world through senses and motor actions. The developmental task of the stage is object permanence - understanding that objects continue to exist when out of sight. Before object permanence develops, an infant loses interest when you hide a toy under a blanket - as if it vanished. That is why peek-a-boo is hilarious to 6-month-olds: the face really does seem to disappear. Object permanence emerges around 8 months.

Also includes the development of stranger anxiety (around 8–12 months), which requires recognizing that the caregiver is distinct from strangers.

Stage 2: Preoperational (2 – 67\frac{6}{7} years)

Children use symbols (words and images) to represent objects and engage in pretend play. But they lack logical “operations” on those symbols.

Key features:

  • Egocentrism. Inability to take another’s perspective. A child hiding their eyes assumes you can’t see them because they can’t see.
  • Magical thinking. Wishes can make things happen; inanimate objects have feelings.
  • Failure of conservation. A child in this stage does not understand that pouring water from a short wide glass into a tall narrow glass leaves the amount unchanged. They judge quantity by appearance.
  • Centration. Focus on one feature at a time (height OR width, not both).

Stage 3: Concrete Operational (7–11 years)

Children master logical reasoning about concrete (real, physical) objects. They can conserve volume, mass, and number. They understand reversibility (the water poured back would be the same amount). They develop the ability to take others’ perspectives (empathy) and perform basic math operations.

Key features:

  • Conservation. Quantity does not change with appearance.
  • Reversibility. Actions can be mentally undone.
  • Decentration. Attention to multiple features at once.
  • Loss of egocentrism.

Stage 4: Formal Operational (12+ years)

Abstract and hypothetical reasoning. Teens can think about things that are not physically present: hypotheticals (“what if gravity didn’t exist”), moral dilemmas, mathematical proofs, and long-term consequences.

Modern developmentalists note that children don’t always hit Piaget’s age targets cleanly, and many adults don’t fully master formal operations across all domains. But the rough progression is widely accepted.

Schemas, Assimilation, and Accommodation

Piaget also gave us the vocabulary of how children (and adults) organize knowledge.

  • Schema. A mental framework that organizes information about a concept. A “dog schema” includes four legs, fur, barks, friendly.
  • Assimilation. Incorporating a new experience into an existing schema. The child sees a wolf and labels it “dog.”
  • Accommodation. Modifying an existing schema to fit a new experience. The child realizes wolves are different and creates a new “wolf” schema.

Cognitive development is a rhythm of assimilation and accommodation, driving toward equilibrium (a consistent set of schemas). Most new information can be assimilated; occasionally something doesn’t fit, and accommodation creates (or splits) a schema.

Problem Solving

Problem solving is moving from a current state to a goal state. Two kinds of problems:

  • Well-defined - clear starting point, clear goal, clear rules. Solving a Sudoku.
  • Ill-defined - ambiguous goal or path. “How do I live a happy life?” or “how do I fit a bookcase around a staircase?”

Methods of Problem Solving

  • Trial and error. Try random solutions until one works. Inefficient; no memory of past attempts.
  • Algorithms. A step-by-step procedure guaranteed to solve the problem. Exhaustive, slow, but reliable. Checking every possible 8-character password is an algorithm.
  • Heuristics. Mental shortcuts that usually (not always) work. Faster than algorithms, but can produce wrong answers.
    • Means-end analysis. Identify the biggest difference between current and goal state, attack that first, repeat.
    • Working backwards. Start from the goal and trace the steps back to the current state. Useful for mazes and math proofs.
  • Intuition. Gut instinct, pattern recognition from experience. Fast but unreliable.
  • Insight. The “aha!” moment when a solution appears seemingly out of nowhere after a period of rest or incubation. Often follows fixation - getting stuck on a bad approach and unable to see alternatives.

Fixation

  • Functional fixedness. Inability to see an object as anything other than its typical use. “I need a nail, but I don’t have a hammer” - a functionally fixed mind doesn’t consider using a wrench or rock to drive the nail.
  • Mental set. A tendency to approach a problem with a strategy that worked before, even when the new situation calls for a different strategy.
At what Piagetian stage is object permanence acquired, and at what age?
Click to reveal answer
Sensorimotor stage (0–2 years). Object permanence emerges around 8 months and is fully developed by the end of the stage.
A 5-year-old insists there is more water in the tall glass than in the short one, even though they saw you pour the same amount. What stage and concept is at play?
Click to reveal answer
Preoperational stage, failure of conservation (the concept that quantity is preserved despite changes in appearance). Conservation develops in the concrete operational stage (around age 7).
Distinguish assimilation from accommodation.
Click to reveal answer
Assimilation: fitting new information into an existing schema ("a wolf is a kind of dog"). Accommodation: modifying or creating schemas to fit information that doesn't fit ("actually wolves are different from dogs - I need a new category").
What is functional fixedness?
Click to reveal answer
A form of cognitive fixation where one cannot see an object as anything other than its conventional use. It blocks creative problem-solving (e.g., not thinking of using a shoe to hammer a nail).
4.5

Heuristics, Biases, and Decision Making

The MCAT loves to ask about heuristics and biases because they are a great way to test whether students understand that human reasoning is systematically, predictably flawed. Most of what follows is the work of Daniel Kahneman and Amos Tversky, who showed that we are not broken rational agents - we are using mental shortcuts that work most of the time and fail in consistent, discoverable ways.

Color photograph of Daniel Kahneman, an older man with gray hair and glasses, looking thoughtfully off-camera
Daniel Kahneman (1934–2024). With Amos Tversky he founded the heuristics-and-biases research program; his work on prospect theory earned the 2002 Nobel Prize in Economics. Credit: nrkbeta via Wikimedia Commons (CC BY-SA 2.0).

Availability Heuristic

Judging the probability of an event by how easily examples come to mind.

  • You hear three shark attack stories on the news this month. You avoid the beach. Actual risk is microscopic, but the recent vivid memories make shark attacks feel more common.
  • Americans consistently overestimate their risk of terrorism and underestimate their risk of heart disease. Terrorism is memorable; heart disease is boring.

The availability heuristic is efficient because frequent events actually do tend to be more memorable. It misfires when vividness, recency, or media coverage make rare events disproportionately available.

Representativeness Heuristic

Judging probability by how well something matches a prototype.

Linda problem (Tversky and Kahneman, 1983). Linda is 31, outspoken, bright, majored in philosophy, and was active in antinuclear demonstrations. Which is more likely: (a) Linda is a bank teller, or (b) Linda is a feminist bank teller?

Most people pick (b). They are wrong. The set of feminist bank tellers is a subset of bank tellers, so (a) must be at least as probable. People fall for (b) because the description is more representative of a feminist than of a bank teller.

This error is called the conjunction fallacy: thinking the probability of two things (A and B) is higher than the probability of one thing (A) alone.

Anchoring-and-Adjustment

Starting from an initial estimate (anchor) and adjusting toward a final answer. The adjustment is usually insufficient, so final answers are biased toward the anchor.

Classic demonstration: ask people whether Gandhi died before or after age 9, and whether he died before or after age 140. Both groups then estimate his actual age at death. The “age 9” group guesses much lower than the “age 140” group, even though both anchors are absurd.

Negotiators use anchoring deliberately. The first price in a negotiation powerfully shapes the final deal, which is why car dealers love to start high.

Biases in Decision Making

  • Overconfidence. People overestimate the accuracy of their own judgments. Why students can walk out of an exam thinking they aced it and score poorly.
  • Belief perseverance. Clinging to a belief despite clear contradicting evidence. Once we form an opinion, new data is easier to ignore than to reckon with.
  • Confirmation bias. Actively seeking out and weighting information that supports existing beliefs while ignoring or discounting contrary information. Why people who believe the earth is flat find “evidence” everywhere.

Framing Effects

The same information, presented differently, produces different decisions.

Tversky and Kahneman’s Asian disease problem: A disease is expected to kill 600 people. Subjects choose between two treatments.

  • Group 1, positive frame:
    • Treatment A: “200 people will be saved.”
    • Treatment B: ”13\frac{1}{3} chance all 600 will be saved, 23\frac{2}{3} chance no one will be saved.”
    • Majority picks A (risk-averse with gains).
  • Group 2, negative frame:
    • Treatment A: “400 people will die.”
    • Treatment B: ”13\frac{1}{3} chance no one will die, 23\frac{2}{3} chance 600 will die.”
    • Majority picks B (risk-seeking with losses).

The options are mathematically identical across groups. The framing alone flips the choice. This is the central insight of prospect theory: people are more sensitive to losses than to equivalent gains (loss aversion), and they take bigger risks to avoid losses than to secure gains.

Prospect theory value function plotted on x-axis (loss to gain in dollars) vs y-axis (subjective value). The curve is S-shaped: in the gain region it rises concavely (diminishing satisfaction with each additional dollar gained); in the loss region it falls convexly and steeply, with the loss side dropping much further than the gain side rises for an equivalent dollar amount
The prospect theory value function. The curve is steeper on the loss side than on the gain side — losing \$5 hurts more than gaining \$5 helps. This loss aversion is why framing the same outcome as a loss flips people's choices toward riskier options. Credit: Laurenrosenberger via Wikimedia Commons (CC BY-SA 4.0).
Define the availability heuristic with an example.
Click to reveal answer
Judging probability by how easily examples come to mind. Example: after watching news about plane crashes, overestimating the risk of flying. Driving to the airport is statistically far riskier.
What is the conjunction fallacy?
Click to reveal answer
Believing that the probability of two events (A AND B) is higher than the probability of one (A alone). Mathematically impossible - A&B is a subset of A. Demonstrated by the Linda/feminist bank teller problem.
A car dealer starts negotiation at \$45,000 for a car actually worth \$28,000. By the end you're paying \$34,000 and feeling good about it. What bias has the dealer exploited?
Click to reveal answer
Anchoring. The high starting price serves as an anchor; your adjustment downward is insufficient, leaving the final price biased toward the anchor. Experienced negotiators set aggressive first offers for exactly this reason.
What does the framing effect demonstrate about human decision making?
Click to reveal answer
Mathematically identical options produce different choices depending on how they are described. People are risk-averse when options are framed as gains and risk-seeking when framed as losses. Central to prospect theory and Kahneman's Nobel Prize work.
4.6

Intelligence

What is intelligence, and do we all have the same kind? Psychologists have been arguing about this for a century. For the MCAT, know the four major theories, the key theorists, and the basic facts about IQ.

One Intelligence or Many?

Spearman’s g (General Intelligence)

Charles Spearman noticed that people who scored well on one kind of mental test tended to score well on others - verbal tests, math tests, spatial tests. He argued for a single underlying factor he called g (general intelligence). Supported by factor analysis: scores across different tests show a common dimension.

If you buy Spearman, IQ tests do capture something real. And they do correlate with many life outcomes (academic success, job performance, health). But g alone doesn’t explain variation in specific abilities.

Thurstone’s Primary Mental Abilities

L.L. Thurstone argued for 7 independent primary mental abilities: word fluency, verbal comprehension, spatial reasoning, perceptual speed, numerical ability, inductive reasoning, and memory. He did not accept that a single g factor could explain everything.

Gardner’s Multiple Intelligences

Howard Gardner expanded the idea further: there are 8–9 independent intelligences, including:

  • Logical-mathematical
  • Verbal-linguistic
  • Spatial-visual
  • Bodily-kinesthetic
  • Interpersonal (understanding others)
  • Intrapersonal (understanding yourself)
  • Musical
  • Naturalist
  • Existential (added later)

Gardner’s theory is popular in educational circles but controversial among psychometricians - the “intelligences” tend to correlate, suggesting they may be variations on g rather than independent abilities.

Sternberg’s Triarchic Theory

Robert Sternberg proposed three intelligences based on real-world success:

  • Analytical intelligence - academic problem-solving (what IQ tests measure).
  • Creative intelligence - novel problem-solving, generating new ideas.
  • Practical intelligence - everyday “street smarts,” adapting to environments.

Fluid vs. Crystallized Intelligence

Cattell distinguished:

  • Fluid intelligence. Ability to reason quickly and abstractly in novel situations. Solving logic puzzles you’ve never seen. Declines with age, starting around 20–30.
  • Crystallized intelligence. Accumulated knowledge and verbal skills. Vocabulary, general knowledge, experience. Stable or improves through middle age, declining slowly only in advanced old age.

Older adults often appear sharp because their crystallized intelligence compensates for modest declines in fluid intelligence. Crossword champions tend to be older.

Emotional Intelligence (EQ)

Proposed by Salovey, Mayer, and popularized by Daniel Goleman. The ability to perceive, understand, manage, and use emotions in self and others. Includes empathy, emotional regulation, and interpersonal effectiveness. Correlates with workplace success and relationship quality, though psychometrically less well validated than g.

IQ Testing: Brief History

  • Alfred Binet (early 1900s, France) developed the first IQ test to identify French children who needed extra help in school. Introduced the idea of mental age - the age at which a child’s intellectual performance is typical.
  • Lewis Terman (Stanford) adapted Binet’s test for American children and adults, creating the Stanford-Binet Intelligence Scale.
  • David Wechsler developed the WAIS (adults) and WISC (children), now the most widely used IQ tests.
  • IQ is normalized so the population mean is 100 with a standard deviation of 15. ~68% of people fall between 85 and 115.
A bell-shaped normal distribution with IQ scores on the x-axis from 55 to 145 and the percentage of population in each band labeled: 0.1% below 55 and above 145, 2.1% in 55-70 and 130-145, 13.6% in 70-85 and 115-130, and 34.1% in each half of 85-115
The IQ bell curve. By construction the mean is 100 and the standard deviation is 15. About 68% of people fall within one SD (85–115), 95% within two SDs (70–130), and 99.7% within three SDs. Credit: Dmcq via Wikimedia Commons (CC BY-SA 3.0).

Nature vs. Nurture for Intelligence

Twin and adoption studies show IQ has a substantial genetic component (~50% heritability in adults) but also a strong environmental component. Identical twins raised apart show higher IQ correlation than fraternal twins raised together, supporting heritability. But environments that deprive children of stimulation, language, or nutrition clearly impair cognitive development.

Modern view: neither genes nor environment alone determines intelligence. Genes set a range; environment determines where in that range an individual lands.

Fixed vs. Growth Mindset

Carol Dweck identified two belief systems about intelligence:

  • Fixed mindset. Intelligence is innate and unchangeable. “Either you’re smart or you’re not.”
  • Growth mindset. Intelligence is malleable with effort and learning. “I can get better at this.”

Students with a growth mindset achieve more over time because they persist through challenges rather than giving up. Praise that emphasizes effort (“you worked hard on this”) fosters a growth mindset; praise that emphasizes ability (“you’re so smart”) can reinforce a fixed mindset.

What is Spearman's g factor?
Click to reveal answer
A single general intelligence factor underlying performance across all mental tasks. Identified through factor analysis showing that scores on different cognitive tests tend to correlate.
Name Sternberg's three intelligences.
Click to reveal answer
Analytical (academic problem solving), creative (generating novel ideas), and practical (everyday "street smarts"). Proposed in his triarchic theory.
Which type of intelligence declines with age, and which stays stable or improves?
Click to reveal answer
Fluid intelligence (abstract reasoning on novel problems) declines starting in early adulthood. Crystallized intelligence (knowledge, vocabulary, experience) stays stable or improves through middle age.
How is IQ standardized?
Click to reveal answer
Population mean = 100, standard deviation = 15. About 68% of people fall within 85–115 (one SD). About 95% fall within 70–130 (two SD).
4.7

Language and the Brain

Language is an astonishing achievement. Children master grammar by age 3 without formal teaching. Adults can process rapid speech, recognize thousands of words, and understand utterances they’ve never heard before. This section covers the brain machinery behind language, the competing theories of how children acquire it, and the debate over whether language shapes thought.

Broca’s Area and Wernicke’s Area

In ~90% of people (including most left-handers), language lives primarily in the left hemisphere. Two key regions:

  • Broca’s area - frontal lobe, inferior frontal gyrus. Responsible for speech production and grammar.
  • Wernicke’s area - temporal lobe, superior temporal gyrus. Responsible for language comprehension.

The two are connected by a bundle of fibers called the arcuate fasciculus, which lets comprehension (Wernicke’s) feed production (Broca’s).

Left-side view of the human brain with Broca's area labeled in the inferior frontal gyrus (front) and Wernicke's area labeled in the posterior superior temporal gyrus (back), with anatomical orientation indicators
The two classical language areas on the dominant (usually left) hemisphere. Broca's area in the frontal lobe drives speech production; Wernicke's area in the posterior temporal lobe handles language comprehension. The arcuate fasciculus (not shown) wires them together. Credit: UX Stalin via Wikimedia Commons (CC BY-SA 4.0).

The Aphasias

Aphasia = a disorder of language production or comprehension. Each region’s damage produces a characteristic pattern.

Broca’s Aphasia (Nonfluent / Expressive)

Damage to Broca’s area. Speech is labored, halted, grammatically simplified. “Water… glass… please.” Comprehension is largely preserved - the patient understands what you say but struggles to produce a fluent response. Patients are often acutely aware and frustrated by their deficit.

Mnemonic: Broca’s is broken speech.

Wernicke’s Aphasia (Fluent / Receptive)

Damage to Wernicke’s area. Speech is fluent and grammatically normal-sounding but nonsensical - “word salad.” Patients may produce made-up words, irrelevant substitutions, and sentences that lack meaning. Comprehension is impaired; the patient does not understand what others say and often does not notice their own errors.

Mnemonic: Wernicke’s comprehension is crappy.

Global Aphasia

Damage to both Broca’s and Wernicke’s areas. Both production and comprehension severely impaired.

Conduction Aphasia

Damage to the arcuate fasciculus (the wiring between Broca’s and Wernicke’s). Production and comprehension are relatively intact, but patients cannot repeat what they just heard. The auditory input reaches comprehension but cannot be routed to the production system.

Split-Brain Patients

The corpus callosum is a thick band of fibers connecting the left and right hemispheres. In rare surgical cases (historically for intractable epilepsy), it is severed to prevent seizures from spreading between hemispheres. The resulting split-brain patients gave neuroscience its richest window into hemispheric specialization.

  • Show a picture to the right visual field (left hemisphere), and the patient can name it - language processing is on the same side.
  • Show the same picture to the left visual field (right hemisphere), and the patient cannot name it - the right hemisphere sees it, but cannot ship the information to the language areas on the left. The patient can pick up the object with their left hand (right hemisphere controls left hand), but cannot say what it is.

Classic split-brain research (Roger Sperry, Nobel 1981) showed that each hemisphere has its own perceptual, cognitive, and even emotional life when disconnected.

Line drawing of a split-brain experimental setup. A patient sits at a table behind a vertical screen, fixating straight ahead. Test objects are placed in front of them under a cloth. A second figure to the right operates a slide projector that flashes images to one visual half-field, beyond the patient's view of their own hands
The classic Sperry split-brain experimental setup. Stimuli are flashed to one visual hemifield while the patient's hand searches a hidden tray of objects — the apparatus needed to test what each disconnected hemisphere can do on its own. Credit: TilmanATuos via Wikibooks/Wikimedia Commons (Public Domain).

Prosody and the Right Hemisphere

While grammar and word meaning live mostly on the left, the right hemisphere contributes prosody - the music of speech: intonation, stress, emotional tone, sarcasm, emphasis. Right-hemisphere damage often spares word-level language but flattens emotional expression and the ability to detect sarcasm or irony in others.

Theories of Language Acquisition

How do children acquire language? Three competing perspectives.

Nativist (Biological) - Noam Chomsky

Color portrait of Noam Chomsky in his late eighties, white hair and beard, looking directly at the camera against a dark background
Noam Chomsky (b. 1928), the linguist whose nativist theory of language acquisition — the language acquisition device and universal grammar — overturned Skinner's behaviorist account in the 1950s. Credit: Σ via Wikimedia Commons (CC BY-SA 4.0).

Children are born with a language acquisition device (LAD) - an innate, specialized mental module for learning language. All languages share a universal grammar at some deep level. The LAD gets activated by exposure and tunes itself to whichever specific language the child hears.

Evidence: children acquire grammar far faster than any general-purpose learning system could explain; they produce grammatically novel sentences they have never heard; the timing of acquisition is strikingly uniform across languages and cultures.

Learning (Behaviorist) - B.F. Skinner

Language is acquired like any other behavior: imitation, reinforcement, and operant shaping. Parents reward correct grammar and approximations; children refine their speech through feedback.

Critique: this cannot easily explain the speed and systematicity of language acquisition or the rarity of explicit correction in most parental speech.

Interactionist

Language acquisition depends on both biological predisposition AND social interaction. Children need contact with fluent speakers at critical periods; language emerges from the interplay of an innate apparatus and meaningful conversation.

The Critical Period Hypothesis

Critical period (or sensitive period) - a window during development (roughly birth to age 8 or so) during which language acquisition is easiest. After the critical period closes, learning a new language becomes significantly harder, particularly grammar and accent. Tragic case studies (like Genie, an abused child isolated from language until age 13) support the existence of a critical period - Genie never fully acquired grammatical language despite extensive teaching.

Does Language Shape Thought?

Universalism

Thought determines language, not the other way around. A tribe with only two words for color can still think about colors; they just don’t have labels.

Linguistic Relativity (Weak Whorfian)

Language influences thought by making some thoughts easier or more habitual, without determining them.

Linguistic Determinism (Strong Whorfian / Sapir-Whorf Hypothesis)

Language determines thought. You cannot think concepts your language doesn’t have words for. The Hopi language is often cited as having no grammatical tense, supposedly preventing Hopi speakers from thinking about time the way English speakers do. Modern linguistics largely rejects the strong version; the weak version has some empirical support.

Piaget and Vygotsky on Language

  • Piaget. Thought develops first; language labels existing cognitive structures. Babies develop object permanence before they learn words like “gone.”
  • Vygotsky. Language and thought develop independently but converge through social interaction. Children internalize the language they hear in conversation, which then becomes the tool of their inner thought.
A patient produces fluent but meaningless speech and cannot understand spoken language. What is the diagnosis and the damaged area?
Click to reveal answer
Wernicke's aphasia. Damage to Wernicke's area in the superior temporal gyrus of the left hemisphere. Speech is fluent and grammatically normal-sounding but semantically incoherent.
What is Chomsky's Language Acquisition Device (LAD)?
Click to reveal answer
An innate, specialized mental mechanism that enables children to acquire any human language rapidly, given exposure during a critical period. Part of Chomsky's nativist theory, which posits a universal grammar underlying all languages.
Differentiate linguistic determinism from linguistic relativism.
Click to reveal answer
Linguistic determinism (strong Whorfian): language DETERMINES what you can think. Linguistic relativism (weak Whorfian): language INFLUENCES thought by making some patterns easier or more habitual. Most modern linguists reject determinism but accept some relativism.
Why might a split-brain patient see a picture, not be able to name it, but still be able to pick it up with their left hand?
Click to reveal answer
If the picture is in the left visual field, it is processed by the right hemisphere. With a severed corpus callosum, the information cannot reach the left hemisphere's language areas, so naming fails. But the right hemisphere still controls the left hand, so picking up works.