Operant Conditioning
Classical conditioning is about reflexes triggered by paired stimuli. Operant conditioning is about voluntary behavior shaped by consequences. B.F. Skinner built the framework: a behavior followed by a good consequence gets repeated; a behavior followed by a bad consequence tapers off.
The Four Quadrants
The single most testable concept in this entire chapter is the 2 × 2 matrix of reinforcement and punishment. Memorize it until you can redraw it in ten seconds.
| Adding something | Removing something | |
|---|---|---|
| Increase behavior | Positive reinforcement | Negative reinforcement |
| Decrease behavior | Positive punishment | Negative punishment |
- “Positive” means adding something. It does NOT mean “good.”
- “Negative” means removing something. It does NOT mean “bad.”
- “Reinforcement” always increases the target behavior.
- “Punishment” always decreases the target behavior.
Examples (the MCAT uses scenarios, so practice classifying them):
- Positive reinforcement. Give a gas gift card to employees for safe driving. You add a reward to increase safe driving.
- Negative reinforcement. Your car’s seatbelt buzzer stops the moment you buckle up. You remove an annoyance to increase seatbelt use.
- Positive punishment. A speeding ticket. You add a financial penalty to decrease speeding.
- Negative punishment. Taking away a teen’s phone for breaking curfew. You remove a valued object to decrease curfew-breaking.
Primary and Secondary Reinforcers
- Primary reinforcers are innately satisfying: food, water, warmth, sex. Useful to animals and infants.
- Secondary reinforcers acquire their value through pairing with primary reinforcers. Money is the classic example. Money has no intrinsic value; it works because we learn it can be traded for primary reinforcers.
Token economies are formal systems built on secondary reinforcers. Patients in a psychiatric unit might earn plastic tokens for doing chores or attending therapy; tokens are later exchanged for privileges. Widely used in schools, prisons, and rehab programs.
Immediacy Matters
Reinforcement and punishment work best when delivered immediately after the target behavior. Delayed consequences weaken the association. This is one reason training a dog works best with real-time treats, and why tax penalties (delivered months or years after the behavior) are a surprisingly weak deterrent. The brain learns “what happened right before” caused the outcome - wait too long and the link blurs.
Escape and Avoidance Learning
Both are forms of aversive control - behavior motivated by the threat of something unpleasant. Both are cases of negative reinforcement.
- Escape learning. The aversive stimulus is already happening; the organism learns a response that terminates it. A rat learns to jump off an electrified grid to end the shock. “Get me out of here.”
- Avoidance learning. A signal precedes the aversive stimulus; the organism learns a response that prevents it from occurring. A warning buzzer sounds before the shock; the rat jumps the barrier at the buzzer and avoids the shock entirely.
Avoidance learning is notoriously durable because the behavior prevents the aversive stimulus from ever happening, which means the learner never finds out whether the stimulus is still present. Phobias have this flavor: someone who avoids flying because of fear never gets to experience a safe flight that would extinguish the fear.
Operant Extinction
If a learned operant behavior stops being reinforced, it gradually stops. A dog trained to sit for treats will stop sitting on command if you never treat again. Like classical extinction, operant extinction is often preceded by an extinction burst - the behavior spikes briefly before fading.
Instinctive Drift
Even well-trained operant behaviors can be overridden by the animal’s species-specific instincts. Instinctive drift, discovered by Breland and Breland (former students of Skinner), is the tendency of trained behaviors to revert to instinctive patterns. They tried to teach a raccoon to deposit tokens in a piggy bank; it started rubbing the tokens together and dunking them - food-washing behavior, not token-depositing. Reinforcement couldn’t override instinct.