Gestalt Principles and Top-Down Processing

Gestalt Principles and Top-Down Processing

6 min read Updated Apr 19, 2026

Your retina catches photons. Your cochlea catches pressure waves. But “seeing a dog” and “hearing music” are not retinal or cochlear events - they are cortical ones. This section covers how your brain assembles sensory signals into a unified perception and why two people can look at the same drawing and see two different images.

Bottom-Up vs. Top-Down Processing

Bottom-up processing starts with the stimulus. Photons hit the retina, signals propagate upward through bipolar cells, ganglion cells, LGN, and visual cortex. Each step adds information. You “build” the percept from raw features. No prior expectation required. This is how a newborn, seeing a face for the very first time, figures out it is a face.

Top-down processing starts with your brain’s expectations. Your brain has a model of what is likely to be out there, and it uses that model to guide interpretation. “Where’s Waldo?” is top-down: you know what Waldo looks like before you start scanning, so your eyes lock onto red-and-white-striped patches. Without the expectation, you would see a chaotic crowd.

The MCAT loves to test one trap: top-down processing can make you perceive things that are not there. When you look at a Kanizsa triangle - three Pac-Man shapes arranged so their mouths form a triangle - you see a white triangle floating in front. There is no triangle. Your brain drew it based on expectation. Top-down is creative; bottom-up is data-driven.

Kanizsa triangle illusion: three black Pac-Man shapes with mouths oriented to suggest the three corners of an equilateral triangle, with three V-shaped line segments between them, making the viewer perceive an illusory white triangle in the center
The Kanizsa triangle. There is no white triangle in the image, yet most viewers see one floating in front. The brain closes the gaps between the Pac-Man cutouts using top-down expectation — a textbook demonstration of illusory contours. Credit: Fibonacci via Wikimedia Commons, CC BY-SA 3.0.

Gestalt Principles: The Whole Is More Than the Sum of the Parts

Gestalt (“form” in German) psychologists argued that our brains group sensory elements into coherent wholes using a handful of rules. Know each one by name, the MCAT will ask.

  • Proximity. Objects near each other are grouped together. Rows of dots spaced closer vertically than horizontally look like columns, not rows.
  • Similarity. Objects that look alike (same shape, color, size) are grouped. A mix of red and blue dots forms two groups even when they are interspersed.
  • Continuity (good continuation). Lines are perceived as following the smoothest path. An X is seen as two crossing lines, not four separate corners.
  • Closure. Incomplete figures are mentally completed. Three Pac-Men suggest a triangle; a broken circle looks like a whole circle.
  • Symmetry. The mind prefers symmetrical groupings, seeing two brackets facing each other as a pair.
  • Figure-ground. Any scene is split into a figure (the object of attention) and a ground (the background). The classic vase/faces illusion is a figure-ground flip.
Rubin's vase figure-ground illusion: a bi-stable image that can be perceived as either a light-colored vase on a black background or two facing silhouetted profiles on a light background
Rubin's vase. Depending on which region your brain assigns as the figure and which as the ground, you see either a vase or two facing profiles — but never both at once. Credit: Nevit Dilmen via Wikimedia Commons, CC BY-SA 3.0.
- **Common fate.** Things moving together are grouped together. A flock of birds looks like a single swirling unit rather than dozens of separate birds. - **Pragnanz (good form / simplicity).** Reality is organized into the simplest form possible. The Olympic rings are seen as five interlocked circles, not a dozen odd curved shapes. This is the overarching principle - every other Gestalt rule is a special case of pragnanz. - **Past experience.** Prior exposure biases grouping. Reading "L" "I" next to each other as two letters rather than the uppercase "U" they could form.
Three Gestalt principles side by side: proximity (dots clumped into groups read as clusters not a uniform grid), similarity (alternating rows of filled and unfilled dots read as horizontal bands), closure (broken outlines of a circle and a rectangle read as complete shapes)
Three of the most testable Gestalt principles. Proximity: dots spaced closer together are grouped as one cluster. Similarity: identical-looking dots form rows even when they interleave. Closure: the brain fills in gaps to complete familiar shapes. Credit: Assembled from Wikimedia Commons files by Kasufcgslfguhvsne et al. (Public Domain).

Depth Perception: Binocular and Monocular Cues

How does a 2D retinal image become a 3D percept? Your brain uses two kinds of cues.

Binocular cues require two eyes.

  • Retinal disparity. Your eyes are about 2.5 inches apart, so they see slightly different images. The bigger the disparity, the closer the object. This is how 3D movies work - they deliver different images to each eye to force disparity.
  • Convergence. When an object is close, your eyes turn inward (eye muscles contract). For a far object, the muscles relax. The brain uses the muscle signal as a distance cue.

Monocular cues work with only one eye.

  • Relative size. If two objects are the same known size but one looks bigger, it must be closer.
  • Interposition (overlap). An object that covers another is in front of it.
  • Relative height. Things higher in the visual field look farther away.
  • Shading and contour. Highlights and shadows suggest 3D form (crater vs. mountain illusion).
  • Motion parallax. When you move, near objects seem to move fast while distant objects drift slowly. Watch power poles streak past the train window while a distant mountain barely budges.
  • Linear perspective. Parallel lines converge at the horizon.
  • Texture gradient. Fine-grained texture looks close; blurred texture looks far.

Perceptual Constancy

Your retinal image changes constantly, but your perception of objects does not. This is perceptual constancy.

  • Size constancy. A friend walking toward you grows on your retina, but you perceive them as the same size.
  • Shape constancy. An opening door casts an increasingly flat trapezoid on your retina, but you still perceive it as rectangular.
  • Color constancy. A white paper looks white under blue sky and warm incandescent bulb even though the light reflected off it differs dramatically. Your brain discounts the illuminant.

Without constancy, objects would seem to change size, shape, and color every time you turned your head. Constancy is a top-down correction that keeps the world stable.

Sensory Adaptation (One More Time)

We covered it briefly in section 1.1, but it is worth closing the chapter on: sensory adaptation is the down-regulation of receptor firing to a constant, unchanging stimulus. You stop feeling your socks after a minute. You stop smelling your own perfume after an hour. You stop hearing a humming air conditioner. The stimulus is still there; your receptors have simply muted themselves so the brain can focus on changes.

Adaptation is why detecting a new stimulus matters more than detecting a steady one. A predator sneaking up behind you is new - motion. A rock behind you is steady and irrelevant.

What is the difference between bottom-up and top-down processing?
Click to reveal answer
Bottom-up starts with raw stimulus data and builds upward. Top-down starts with expectations and interprets ambiguous data through prior knowledge. Top-down can generate percepts that don't exist (illusions).
Name three binocular depth cues.
Click to reveal answer
Retinal disparity (slight difference between images in the two eyes) and convergence (inward rotation of the eyes on near objects). Those are the two primary ones; most "depth cues" lists have only these two as binocular, with everything else monocular.
What is pragnanz (the law of good form)?
Click to reveal answer
The Gestalt principle that the brain perceives ambiguous or complex images in the simplest way possible. The Olympic rings look like five circles rather than a dozen separate curves.
What is perceptual constancy?
Click to reveal answer
The brain perceives objects as having stable properties (size, shape, color) despite continuous changes in the retinal image. Size, shape, and color constancy are the classic three.