Operant Conditioning

Operant Conditioning

6 min read Updated Apr 19, 2026

Classical conditioning is about reflexes triggered by paired stimuli. Operant conditioning is about voluntary behavior shaped by consequences. B.F. Skinner built the framework: a behavior followed by a good consequence gets repeated; a behavior followed by a bad consequence tapers off.

Labeled schematic of an operant conditioning chamber (Skinner box) with a rat inside, showing loudspeakers and lights at the top as discriminative stimuli, a response lever in the middle, a food dispenser on one side, and an electrified grid floor for aversive stimuli
The operant conditioning chamber ("Skinner box"). The animal emits a response (pressing the lever); the apparatus delivers a consequence (food pellet or shock). Varying the schedule and contingency of the consequence is how Skinner built the entire experimental science of operant conditioning. Credit: AndreasJS via Wikimedia Commons, CC BY-SA 3.0.

The Four Quadrants

The single most testable concept in this entire chapter is the 2 × 2 matrix of reinforcement and punishment. Memorize it until you can redraw it in ten seconds.

Adding somethingRemoving something
Increase behaviorPositive reinforcementNegative reinforcement
Decrease behaviorPositive punishmentNegative punishment
  • “Positive” means adding something. It does NOT mean “good.”
  • “Negative” means removing something. It does NOT mean “bad.”
  • “Reinforcement” always increases the target behavior.
  • “Punishment” always decreases the target behavior.

Examples (the MCAT uses scenarios, so practice classifying them):

  • Positive reinforcement. Give a gas gift card to employees for safe driving. You add a reward to increase safe driving.
  • Negative reinforcement. Your car’s seatbelt buzzer stops the moment you buckle up. You remove an annoyance to increase seatbelt use.
  • Positive punishment. A speeding ticket. You add a financial penalty to decrease speeding.
  • Negative punishment. Taking away a teen’s phone for breaking curfew. You remove a valued object to decrease curfew-breaking.

Primary and Secondary Reinforcers

  • Primary reinforcers are innately satisfying: food, water, warmth, sex. Useful to animals and infants.
  • Secondary reinforcers acquire their value through pairing with primary reinforcers. Money is the classic example. Money has no intrinsic value; it works because we learn it can be traded for primary reinforcers.

Token economies are formal systems built on secondary reinforcers. Patients in a psychiatric unit might earn plastic tokens for doing chores or attending therapy; tokens are later exchanged for privileges. Widely used in schools, prisons, and rehab programs.

Immediacy Matters

Reinforcement and punishment work best when delivered immediately after the target behavior. Delayed consequences weaken the association. This is one reason training a dog works best with real-time treats, and why tax penalties (delivered months or years after the behavior) are a surprisingly weak deterrent. The brain learns “what happened right before” caused the outcome - wait too long and the link blurs.

Escape and Avoidance Learning

Both are forms of aversive control - behavior motivated by the threat of something unpleasant. Both are cases of negative reinforcement.

  • Escape learning. The aversive stimulus is already happening; the organism learns a response that terminates it. A rat learns to jump off an electrified grid to end the shock. “Get me out of here.”
  • Avoidance learning. A signal precedes the aversive stimulus; the organism learns a response that prevents it from occurring. A warning buzzer sounds before the shock; the rat jumps the barrier at the buzzer and avoids the shock entirely.

Avoidance learning is notoriously durable because the behavior prevents the aversive stimulus from ever happening, which means the learner never finds out whether the stimulus is still present. Phobias have this flavor: someone who avoids flying because of fear never gets to experience a safe flight that would extinguish the fear.

Operant Extinction

If a learned operant behavior stops being reinforced, it gradually stops. A dog trained to sit for treats will stop sitting on command if you never treat again. Like classical extinction, operant extinction is often preceded by an extinction burst - the behavior spikes briefly before fading.

Instinctive Drift

Even well-trained operant behaviors can be overridden by the animal’s species-specific instincts. Instinctive drift, discovered by Breland and Breland (former students of Skinner), is the tendency of trained behaviors to revert to instinctive patterns. They tried to teach a raccoon to deposit tokens in a piggy bank; it started rubbing the tokens together and dunking them - food-washing behavior, not token-depositing. Reinforcement couldn’t override instinct.

A parent takes away their teen's video games when the teen misses chores. Classify this operant technique.
Click to reveal answer
Negative punishment. Something desirable (video games) is removed, to decrease a behavior (missing chores).
What is the difference between negative reinforcement and punishment?
Click to reveal answer
Negative reinforcement INCREASES a behavior by removing something unpleasant (e.g., buckle seatbelt → buzzer stops). Punishment DECREASES a behavior. They are opposites, not synonyms.
Why is avoidance learning so resistant to extinction?
Click to reveal answer
The avoidance response prevents the aversive stimulus, so the learner never discovers whether the stimulus is still present. The behavior is self-reinforcing via relief. Phobic avoidance works this way.
What is instinctive drift?
Click to reveal answer
The tendency of operantly trained behavior to revert to species-specific instinctive behavior. The Brelands showed raccoons reverting to food-washing motions instead of token-depositing, despite reinforcement. Instincts can override learned behavior.