Schedules of Reinforcement and Shaping
Skinner discovered that when you reward behavior is as important as whether you reward it. Different reinforcement schedules produce dramatically different response patterns and different resistance to extinction. This is why slot machines keep people pulling levers and factory pieceworkers keep producing. Understanding the four partial schedules is worth real MCAT points.
Continuous vs. Partial Reinforcement
- Continuous reinforcement. Every instance of the target behavior is reinforced. A vending machine: press the button, get the snack, every time. Continuous reinforcement produces fast initial learning but also fast extinction - if the machine suddenly stops working, you quickly stop pressing.
- Partial (intermittent) reinforcement. Only some instances of the behavior are reinforced. Partial schedules produce slower learning but much greater resistance to extinction. Because the learner has already experienced non-reward, they keep going even when rewards stop.
The resistance-to-extinction difference is why occasional reinforcement is often more addictive than constant reinforcement. A gambler who wins every time on a slot machine quits the moment the machine breaks. A gambler who wins sometimes keeps pulling.
The Four Partial Schedules
Two axes: ratio or interval (by behavior count or by time elapsed), and fixed or variable (constant or unpredictable).
| Fixed | Variable | |
|---|---|---|
| Ratio (count-based) | Fixed ratio (FR) | Variable ratio (VR) |
| Interval (time-based) | Fixed interval (FI) | Variable interval (VI) |
Fixed Ratio (FR)
Reinforcement after a fixed number of responses. Example: a factory worker gets paid for every 10 widgets assembled.
- Produces a high, steady response rate with a brief pause after each reinforcement (the worker rests briefly after collecting their pay).
- Resistance to extinction: moderate.
Variable Ratio (VR)
Reinforcement after an unpredictable number of responses, averaging around some set value. Example: a slot machine pays out on average every 15 pulls, but any specific pull could win or lose.
- Produces the highest, most consistent response rate of any schedule.
- Most resistant to extinction. The learner never knows if the next response will be the jackpot, so quitting feels costly.
- Powers most addictive behavior: gambling, social media (“maybe the next scroll has something good”), fishing, sales cold-calling.
Fixed Interval (FI)
First response after a fixed amount of time earns reinforcement. Example: a salaried employee gets paid every two weeks regardless of effort.
- Produces a “scalloped” response pattern - slow responding just after reinforcement, ramping up as the next reinforcement approaches. Think of a student doing no work for the first week after a test, then cramming as the next test nears.
- Low overall response rate.
Variable Interval (VI)
Reinforcement after an unpredictable amount of time. Example: checking your email - you don’t know when the next important message will arrive.
- Produces a steady, moderate response rate.
- High resistance to extinction.
- Pop quizzes work on this schedule (reinforce studying by unpredictable testing).
Shaping: Building Complex Behaviors
How do you get an animal to do something it would never do spontaneously? Shaping is the reinforcement of successive approximations - rewarding behaviors that get progressively closer to the target behavior.
To teach a pigeon to peck a specific button:
- Reward the pigeon for facing the button.
- Once it reliably faces the button, require it to step toward the button before rewarding.
- Then require it to touch the button.
- Finally, reward only when it pecks the button.
Each step raises the bar slightly. The pigeon learns a complex chain of behavior it could never have guessed from reading the original instructions. Trainers use shaping to teach dogs to skateboard, dolphins to jump through hoops, and humans to do surgery.
Chaining
A related concept: chaining links individual operant responses into a sequence. Each response serves as a cue for the next. Dance routines, driving a stick shift, cooking a complex recipe - all are chains of simpler learned responses that have been stitched together through reinforcement.