Amino Acids, Peptides, and Proteins

Chapter 1: Amino Acids, Peptides, and Proteins

Full chapter view · 12 sections · ~55 min read Switch to section-by-section view →
1.1

Amino Acid Structure

Every amino acid in your body is built on the same four-part chassis. Like a car - wheels, engine, frame, seats - the parts never change. What changes is the paint job. That paint job is the side chain, and it decides whether the amino acid is greasy, charged, acidic, basic, bulky, or tiny.

If you understand one amino acid, you understand the structural skeleton of all 20.

The Four Groups Around the Central Carbon

Every standard amino acid has a central carbon (the alpha carbon, Cα) with four things attached:

  1. An amino group (-NH2 or -NH3+ at physiological pH)
  2. A carboxyl group (-COOH or -COO- at physiological pH)
  3. A hydrogen atom
  4. An R group (the side chain - this is what varies)

The name “amino acid” comes straight from the structure: an “amino” group and a carboxylic “acid” group on the same carbon.

General amino acid structure with the central alpha carbon bonded to an amino group, a carboxyl group, a hydrogen atom, and an R group side chain
The general amino acid structure. Every standard amino acid has this same four-part chassis around the alpha carbon; only the R group changes. Credit: OpenStax Biology 2e, CC BY 4.0

The Alpha Carbon is Chiral (Except in Glycine)

Because the alpha carbon has four different groups attached in 19 of the 20 standard amino acids, it is a chiral center. That means it has two non-superimposable mirror images: L and D forms.

Biological proteins use L-amino acids almost exclusively. Your ribosomes only build with the L-form. D-amino acids do exist (bacterial cell walls use some, and a few show up in venoms and signaling peptides) but they are passage-peek rare on the MCAT.

Glycine is the exception. Its R group is just a hydrogen atom, so the alpha carbon has two identical hydrogens. No chirality. Glycine is achiral.

Why the Structure Matters

The amino and carboxyl groups on the central carbon are what allow amino acids to link together into chains. The carboxyl of one amino acid reacts with the amino of the next, forming a peptide bond and releasing water. That reaction happens millions of times per second on a ribosome. We cover the mechanism in peptide bond formation.

The R group is what makes each amino acid unique. Some R groups are just plain alkyl chains (greasy). Some have polar OH groups. Some have carboxylic acids (negatively charged) or amines (positively charged) built in. Some have aromatic rings. A tiny few do special tricks (like cysteine forming disulfide bridges). We sort all 20 of them in the next section.

What four groups are attached to the alpha carbon of a standard amino acid?
Click to reveal answer
(1) An amino group (-NH2 / -NH3+), (2) a carboxyl group (-COOH / -COO-), (3) a hydrogen atom, and (4) a side chain (R group). The R group is the only one that varies between the 20 standard amino acids.
Which amino acid is achiral, and why?
Click to reveal answer
Glycine. Its R group is a single hydrogen, so the alpha carbon has two identical hydrogens - only three unique substituents. Four different groups are required for chirality. Glycine is the only standard amino acid without a stereocenter.
Which stereoisomer of amino acids is used to build proteins in eukaryotic cells?
Click to reveal answer
The L-form (L-amino acids). Ribosomes read the genetic code and incorporate only L-amino acids. D-amino acids occur in nature (bacterial cell walls, some peptides) but are not standard in eukaryotic proteins.
1.2

Classification

Memorizing 20 random structures is a slog. Grouping them by what their R group does is a cheat code. Once you see that histidine, lysine, and arginine all have nitrogen that grabs a proton and becomes positive, you know what they do in a protein regardless of which one the passage names.

The 20 standard amino acids split into five main buckets plus three special cases. Learn the buckets first. Individual structures come in the next section.

The Five R-Group Buckets

CategoryWhat the R group doesMembersAt pH 7 charge
Nonpolar aliphaticGreasy, hates water, buried inside proteinsGly, Ala, Val, Leu, Ile, Met, Pro0
AromaticHas a benzene-like ring, mostly hydrophobicPhe, Trp, Tyr (Tyr is slightly polar from its -OH)0
Polar unchargedHas -OH, -SH, or amide, loves water but no chargeSer, Thr, Cys, Asn, Gln0
Positively charged (basic)Nitrogen that grabs H+Lys, Arg, His+1 (His often partial)
Negatively charged (acidic)Extra -COOH that gives up H+Asp, Glu-1
Chart of all 20 standard amino acids grouped by R group category: nonpolar aliphatic, nonpolar aromatic, polar uncharged, positively charged, and negatively charged, each with structural formulas labeled
All 20 standard amino acids grouped by R-group property. Color-coded blocks show each side chain highlighted. Learn the groups first, then the individual structures. Credit: OpenStax Biology 2e, CC BY 4.0

Nonpolar R Groups (Hydrophobic)

These have alkyl chains, sulfur atoms (in methionine), or aromatic rings with no polar groups. They are greasy. In a folded protein, they huddle in the interior away from water - the hydrophobic core. Aromatic residues also stack with each other for extra stability.

Polar Uncharged R Groups

These have heteroatoms (O, N, S) that form hydrogen bonds but do not carry a net charge at pH 7. They love water. They sit on the protein surface or in active sites where they can hydrogen bond to substrates.

  • Serine, Threonine - have -OH groups
  • Tyrosine - aromatic ring with -OH (classified polar or aromatic depending on the chart)
  • Cysteine - has -SH (thiol)
  • Asparagine (N), Glutamine (Q) - have amide side chains

Positively Charged R Groups (Basic)

Three amino acids. Their side chains carry a nitrogen lone pair that picks up a proton and becomes positively charged at pH 7.

  • Lysine (K) - long chain ending in -NH3+
  • Arginine (R) - guanidinium group, very basic (pKa about 12, fully protonated at pH 7)
  • Histidine (H) - imidazole ring, pKa near 6, so at pH 7 it is partly protonated. This makes histidine uniquely useful in enzyme active sites that need to shuttle protons.

Negatively Charged R Groups (Acidic)

Two amino acids, both ending in “-ate” because they donate their carboxyl proton at pH 7.

  • Aspartate (Asp, D) - one -CH2- linker, then -COO-
  • Glutamate (Glu, E) - two -CH2- linkers, then -COO-

Their protonated forms are called aspartic acid and glutamic acid, but at physiological pH they are almost always in the deprotonated “-ate” form.

The Three Special Cases

Glycine (Gly, G). R group is just -H. Smallest amino acid. Only achiral one. Flexible - fits anywhere. Often found at tight turns in a protein.

Proline (Pro, P). Its R group loops back and bonds to the amino nitrogen, forming a rigid five-membered ring. This kink disrupts alpha helices and creates helix breakers or sharp bends.

Cysteine (Cys, C). Its -SH (thiol) can form a covalent disulfide bond (-S-S-) with another cysteine. This is the only covalent bond (besides peptide bonds) that stabilizes folded proteins.

Which three amino acids are positively charged at physiological pH?
Click to reveal answer
Lysine (K), Arginine (R), and Histidine (H). Arg is the most basic (pKa about 12). His has a pKa near 6, so it can be protonated or deprotonated at pH 7, which is why histidine is so common in enzyme active sites that shuttle protons.
Why is proline often called a "helix breaker"?
Click to reveal answer
Its R group loops back and covalently bonds to the backbone nitrogen, forming a rigid 5-membered ring. This locked conformation and the absence of a free N-H for hydrogen bonding disrupt the regular pattern of an alpha helix, forcing a kink or ending the helix.
Which amino acid can form covalent disulfide bridges, and why does that matter?
Click to reveal answer
Cysteine, because it has an -SH (thiol) side chain. Two cysteines can be oxidized to form a covalent S-S bond, locking distant parts of a protein together. This is the strongest non-backbone stabilizing force in tertiary structure and is critical in secreted/extracellular proteins like antibodies and insulin.
1.3

The 20 Amino Acids

You need to know all 20 by structure, name, three-letter code, one-letter code, and category. On test day you may see any of these formats in a passage and have to reason about it instantly. The table below is worth memorizing cold.

The twenty amino acids, sorted by what the side chain does

Reference
Ask three questions of any side chain: does it carry charge at pH 7, is it polar, and is it bulky. Nonpolar buried in the core Gly G not chiral Ala A Alanine Val V Valine Leu L Leucine Ile I Isoleucine Pro P kinks the chain Met M starts translation Aromatic absorb at 280 nm Phe F Phenylalanine Trp W absorbs most Tyr Y can be phosphorylated Polar on the surface Ser S phosphorylated Thr T phosphorylated Cys C makes disulfides Asn N Asparagine Gln Q Glutamine Acidic negative at pH 7 Asp D pKa ≈ 3.9 Glu E pKa ≈ 4.3 Basic positive at pH 7 Lys K pKa ≈ 10.5 Arg R pKa ≈ 12.5 His H pKa ≈ 6.0 Essential your body cannot build these, so they have to be eaten His · Ile · Leu · Lys · Met · Phe · Thr · Trp · Val Glycine
Side chain is a single hydrogen, so it has no chiral centre and fits into turns nothing else can.
Proline
Its side chain loops back onto its own nitrogen. No backbone N-H to donate, so it breaks α-helices.
Cysteine
Two of them oxidize into a covalent disulfide bridge, the only side-chain-to-side-chain bond.
1

Scroll sideways to see the whole map.

Acidic: deprotonated and negative at pH 7 Basic: protonated and positive at pH 7 Polar but uncharged Aromatic
You are not asked to draw these; you are asked to predict behavior from them. Read a residue and ask three questions: does it carry charge at pH 7, is it polar, and is it big. Those three answers cover almost every amino acid question the exam sets.

The Master Table

| Name | 3-letter | 1-letter | Category | Side chain feature | pKa of R group (if any) |
|------|----------|----------|----------|--------------------|-------------------------|
| Glycine | Gly | G | Special (nonpolar) | -H (achiral) | - |
| Alanine | Ala | A | Nonpolar aliphatic | -CH3 | - |
| Valine | Val | V | Nonpolar aliphatic (branched) | -CH(CH3)2 | - |
| Leucine | Leu | L | Nonpolar aliphatic (branched) | -CH2CH(CH3)2 | - |
| Isoleucine | Ile | I | Nonpolar aliphatic (branched) | -CH(CH3)CH2CH3 | - |
| Methionine | Met | M | Nonpolar, sulfur | -CH2CH2SCH3 | - |
| Proline | Pro | P | Special (nonpolar) | Cyclic, bonds back to N | - |
| Phenylalanine | Phe | F | Aromatic (nonpolar) | -CH2-C6H5 | - |
| Tryptophan | Trp | W | Aromatic (nonpolar) | Indole ring | - |
| Tyrosine | Tyr | Y | Aromatic (polar) | Phenol ring (-OH) | ~10 |
| Serine | Ser | S | Polar uncharged | -CH2OH | - |
| Threonine | Thr | T | Polar uncharged | -CH(OH)CH3 | - |
| Cysteine | Cys | C | Special (polar) | -CH2SH | ~8.3 |
| Asparagine | Asn | N | Polar uncharged | -CH2CONH2 | - |
| Glutamine | Gln | Q | Polar uncharged | -CH2CH2CONH2 | - |
| Aspartate | Asp | D | Acidic (negative) | -CH2COO- | ~3.9 |
| Glutamate | Glu | E | Acidic (negative) | -CH2CH2COO- | ~4.1 |
| Lysine | Lys | K | Basic (positive) | -(CH2)4NH3+ | ~10.5 |
| Arginine | Arg | R | Basic (positive) | Guanidinium | ~12.5 |
| Histidine | His | H | Basic (partial +) | Imidazole | ~6.0 |

The alpha-amino group of any free amino acid has a pKa around 9 to 10. The alpha-carboxyl group has a pKa around 2. Memorize these two values - they appear in every titration question.

Essential vs. Nonessential

Humans can synthesize 11 of the 20 amino acids from scratch. The other 9 must come from diet - these are the essential amino acids.

Amino Acids with Ionizable Side Chains

Seven amino acids have R groups with ionizable protons. These are the only ones whose charge state depends on pH, and they are the ones you will see on titration and isoelectric-point questions.

| Amino acid | R-group pKa | Charge below pKa | Charge above pKa |
|------------|-------------|------------------|------------------|
| Asp (D) | 3.9 | 0 | -1 |
| Glu (E) | 4.1 | 0 | -1 |
| His (H) | 6.0 | +1 | 0 |
| Cys (C) | 8.3 | 0 | -1 |
| Tyr (Y) | 10.1 | 0 | -1 |
| Lys (K) | 10.5 | +1 | 0 |
| Arg (R) | 12.5 | +1 | 0 |

How to Study This Table

Do not try to memorize everything at once. Start with categories (previous section). Then memorize the three-letter codes by writing them out with structures five times. Then add the one-letter codes. Then pKa values for the seven ionizable ones. This layered approach beats rote flashcards.

Which seven amino acids have ionizable side chains?
Click to reveal answer
Aspartate (D, pKa ~3.9), Glutamate (E, pKa ~4.1), Histidine (H, pKa ~6.0), Cysteine (C, pKa ~8.3), Tyrosine (Y, pKa ~10.1), Lysine (K, pKa ~10.5), Arginine (R, pKa ~12.5). These are the amino acids whose charge state depends on pH and that appear in titration questions.
What are the one-letter codes for the 9 essential amino acids?
Click to reveal answer
F (Phe), V (Val), T (Thr), W (Trp), I (Ile), M (Met), H (His), K (Lys), L (Leu). Mnemonic “PVT TIM HALL” covers them (plus Arginine as conditionally essential).
Which amino acid has a pKa near 6 and is commonly found in enzyme active sites as a proton shuttle?
Click to reveal answer
Histidine. Its imidazole side chain has a pKa of about 6, which means at physiological pH (7.4) both the protonated and deprotonated forms coexist. This lets histidine both donate and accept protons during catalysis - hence its starring role in enzymes like chymotrypsin and carbonic anhydrase.
1.4

Acid-Base Chemistry

Amino acids have two ionizable groups always (the alpha amino and alpha carboxyl) plus one more if the R group is ionizable. At any given pH, some of those groups are protonated and some are not. Combining those charge states across three groups gives every amino acid its personality on a titration curve.

Titration and the isoelectric point

Acid-base
0 2 4 6 8 10 12 14 0 1 2 equivalents of base added pH pKa₁ ≈ 2.3 the –COOH group pKa₂ ≈ 9.6 the –NH₃⁺ group pI ≈ 6.0 net charge zero Cation +1 NH₃⁺ · COOH Zwitterion 0 NH₃⁺ · COO⁻ Anion −1 NH₂ · COO⁻ With a third group an ionizable side chain adds a third plateau and moves pI Acidic side chain Asp, Glu average the two lowest pKa pI ≈ 3 Basic side chain Lys, Arg, His average the two highest pKa pI ≈ 10 Below pI a molecule is net positive. Above pI it is net negative. That one rule runs every separation method.
1

Scroll sideways to see the whole map.

The curve is a charge readout, not just a chemistry plot. Read left to right and the molecule loses protons one group at a time: fully protonated and positive, then neutral and zwitterionic, then fully deprotonated and negative. Everything the exam asks about pI is somewhere on that sentence.

You will see this concept tested two ways: (1) “what is the charge at pH X” and (2) “what is the pI.” Master both and you can handle every acid-base amino acid question.

The Zwitterion

A zwitterion is a molecule with both a positive and a negative charge that cancel out to give a net charge of zero. At physiological pH (7.4), every amino acid with a nonionizable R group exists almost entirely as a zwitterion: the alpha amino group is protonated (-NH3+) and the alpha carboxyl is deprotonated (-COO-).

Titration Curve of a Simple Amino Acid (Glycine)

Glycine has two ionizable groups. Titrating from low pH (excess H+) with added base gives a curve with two plateaus and three charged species.

At very low pH: fully protonated. Both -NH3+ and -COOH. Net charge = +1.

As pH rises past pKa1 (~2.3): the -COOH loses its proton to become -COO-. Now -NH3+ and -COO-. Net charge = 0. This is the zwitterion. The flat region near pKa1 is a buffering zone.

As pH rises past pKa2 (~9.6): the -NH3+ loses its proton to become -NH2. Now -NH2 and -COO-. Net charge = -1.

Calculating the Isoelectric Point (pI)

The pI is the pH at which the amino acid carries no net charge. For an amino acid with only the two backbone ionizable groups, the pI is simply:

With an Ionizable Side Chain

If the R group is also ionizable, the amino acid has three pKa values. The pI is the average of the two pKas that flank the zwitterionic (net-zero) form.

  • Acidic amino acid (Asp, Glu): pI = average of the two lowest pKas (both the side-chain carboxyl and the alpha carboxyl lose protons to produce the neutral form). The pI is low (around 3 to 4).
  • Basic amino acid (Lys, Arg, His): pI = average of the two highest pKas (both amino/basic groups must still be protonated for the zwitterion). The pI is high (around 7.6 to 11).
  • Neutral amino acid with an ionizable side chain (Cys, Tyr): treat it like a neutral amino acid; pI is the average of the alpha-carboxyl pKa and the alpha-amino pKa.

Worked Examples: Three Types, Three Calculations

The MCAT loves these. Work through all three types once and the pattern will feel obvious on test day.

The shortcut: acidic side chain → pI = average of the two lowest pKas; basic side chain → pI = average of the two highest pKas. This works because you always sandwich the zwitterion between the two pKas closest to it on the protonation ladder.

Charge at Any pH

Two simple rules beat every “what is the charge” question:

  1. If pH is below pKa, the group is protonated.
  2. If pH is above pKa, the group is deprotonated.

For an amine (-NH3+ / -NH2): protonated is +1, deprotonated is 0.
For a carboxyl (-COOH / -COO-): protonated is 0, deprotonated is -1.

Apply the rule to every ionizable group on the molecule and sum the charges.

Henderson-Hasselbalch for Amino Acids

The Henderson-Hasselbalch equation applies to each ionizable group:

What is a zwitterion?
Click to reveal answer
A molecule that has both a positive and a negative charge which cancel to give a net charge of zero. For a typical amino acid at physiological pH, the alpha amino is protonated (-NH3+) and the alpha carboxyl is deprotonated (-COO-) simultaneously.
How do you calculate pI for a neutral amino acid (no ionizable side chain)?
Click to reveal answer
pI = (pKa1 + pKa2) / 2, the average of the alpha-carboxyl pKa (~2) and the alpha-amino pKa (~9). For glycine: (2.34 + 9.60) / 2 = 5.97.
How do you find the pI for an acidic amino acid like glutamate (pKa values ~2.2, 4.1, 9.7)?
Click to reveal answer
Average the two pKas that flank the zwitterionic (net-zero) form. For glutamate, both carboxyls (alpha-COOH pKa ~2.2 and side-chain pKa ~4.1) must be deprotonated for net-zero, so pI = (2.2 + 4.1) / 2 ≈ 3.15. Acidic amino acids have low pI values; basic amino acids have high pI values.
1.5

Peptide Bond Formation

Snap LEGO bricks together. That is peptide bond formation. Take the carboxyl group of one amino acid, the amino group of the next, pop off a water molecule, and you have a covalent amide bond linking the two. Do this 146 times in a row and you have the hemoglobin beta chain. Do it 1,480 times and you have a mid-size enzyme.

This is the reaction that builds every protein in your body, at a rate of about 10 to 20 amino acids per second on a ribosome.

The Mechanism

The -COOH of amino acid #1 reacts with the -NH2 of amino acid #2. The -OH of the carboxyl leaves together with one H from the amino group - that is the water molecule released.

What remains is a covalent bond from the carbonyl carbon of amino acid #1 directly to the nitrogen of amino acid #2. The functional group -C(=O)-N(H)- is an amide. A peptide bond is just a specific kind of amide bond that connects two amino acids.

Peptide bond formation diagram showing the carboxyl group of one amino acid reacting with the amino group of another, releasing water, and forming a peptide bond between the carbonyl carbon and the nitrogen atom
Peptide bond formation is a condensation (dehydration synthesis) reaction. The -OH from the carboxyl group and an -H from the amino group leave as water; a covalent C-N bond forms. Credit: OpenStax Biology 2e, CC BY 4.0

Naming the Ends: N-Terminus and C-Terminus

Every polypeptide has two ends because each peptide bond uses up one -NH2 and one -COOH but leaves the amino group of the first residue and the carboxyl of the last residue free.

  • The N-terminus has a free alpha amino group. This is where translation starts.
  • The C-terminus has a free alpha carboxyl group. This is where translation ends.

Peptide sequences are always written N-terminus to C-terminus, left to right. “Ala-Gly-Ser” means alanine at the N-terminus, glycine in the middle, serine at the C-terminus.

Energy Cost

Forming a peptide bond is thermodynamically unfavorable (positive ΔG) under cellular conditions. Your ribosome pays for it: each new amino acid costs 4 high-energy phosphate bonds (1 ATP to charge the tRNA with its amino acid, plus 2 GTP during elongation and 1 GTP during translocation). That is why protein synthesis is among the most energy-expensive processes in the cell.

Breaking Peptide Bonds: Hydrolysis

The reverse reaction breaks a peptide bond by adding water back. Your stomach acid and digestive enzymes (pepsin, trypsin, chymotrypsin) all catalyze peptide bond hydrolysis. Without an enzyme, the reaction is glacially slow - peptide bonds have half-lives of hundreds of years in neutral water. With the right enzyme, it happens in milliseconds.

Specificity: Why Sequencing Works

Not all proteases are created equal. Each one recognizes specific amino acid residues at the cleavage site, which is what makes peptide mapping and sequencing possible. If you already know the cutting rules, you can predict exactly which fragments a protease will generate from a given sequence.

Peptide Length Terminology

  • Dipeptide: 2 amino acids, 1 peptide bond
  • Tripeptide: 3 amino acids, 2 peptide bonds
  • Oligopeptide: a few (usually <10-20) amino acids
  • Polypeptide: longer chain, but not yet folded or functional
  • Protein: one or more polypeptide chains folded into a functional 3D shape

Note that for n amino acids, there are always (n - 1) peptide bonds.

What functional group is lost when a peptide bond forms, and what is the byproduct?
Click to reveal answer
The -OH from the carboxyl of one amino acid and one -H from the amino group of the next are lost together as a water molecule (H2O). This is a condensation (dehydration synthesis) reaction. The covalent C-N bond that forms is an amide bond, specifically called a peptide bond.
A peptide is written as "Gly-Trp-Lys-Asp." Which residue is at the N-terminus?
Click to reveal answer
Glycine. Peptide sequences are always written N-terminus to C-terminus, left to right. So Gly is the N-terminus and Asp is the C-terminus. The ribosome also builds the chain in that order - N first, then C last.
How many peptide bonds are in a protein with 150 amino acids?
Click to reveal answer
149. In general, n amino acids in a single chain are connected by (n - 1) peptide bonds, because each new residue added to the chain forms exactly one new bond.
1.6

Peptide Bond Properties

A peptide bond looks like a single C-N bond. It isn’t. Two electron resonance structures share the electrons between a C-N and a C=N form. The result is partial double bond character, which has three consequences you need to know cold:

  1. The peptide bond is planar (all six atoms in the amide group lie in a flat plane).
  2. It is rigid (no rotation around the peptide bond itself).
  3. The substituents on either side adopt the trans configuration in almost every case.

Everything about how proteins fold starts from these three facts.

Why the Peptide Bond is Planar

The nitrogen in an amide has a lone pair. That lone pair delocalizes into the neighboring carbonyl C=O. The resonance structure puts a double bond between C and N and a negative charge on the oxygen. In the real molecule, this is an average: the C-N bond has about 40 percent double-bond character.

Double bonds cannot rotate. Therefore the peptide bond cannot rotate. All six atoms involved - the alpha carbon of residue i, the carbonyl C and O, the amide N and H, and the alpha carbon of residue i+1 - lie in a single plane.

Why Trans is Preferred

When two alpha carbons are bonded through a planar peptide bond, the R groups can either be on the same side of the plane (cis) or opposite sides (trans). Trans is strongly preferred because it keeps the bulky R groups away from each other. Cis peptide bonds cause steric clashes.

About 99.9 percent of peptide bonds in folded proteins are trans. The rare exception: proline. Because proline’s side chain loops back to the amide nitrogen, the cis and trans forms are closer in energy, and about 5 to 10 percent of X-Pro peptide bonds (peptide bond immediately before proline) are cis.

Where the Backbone Can Rotate: Phi and Psi

If the peptide bond itself is frozen, where does the flexibility come from? The two bonds on either side of each alpha carbon. These are called phi (Φ) and psi (Ψ):

  • Phi (Φ): the rotation around the N-Cα bond.
  • Psi (Ψ): the rotation around the Cα-C(=O) bond.

Every residue in a polypeptide has its own Φ and Ψ. Together, these two angles for every residue define the entire backbone geometry of the protein.

The Ramachandran Plot

Because of steric clashes between the R group and the backbone, not every combination of Φ and Ψ is physically possible. Plotting Φ vs. Ψ for all residues in known protein structures gives the Ramachandran plot. Only a few clusters of allowed angles appear:

  • The alpha helix region (around Φ = -60°, Ψ = -45°)
  • The beta sheet region (around Φ = -120°, Ψ = +120°)
  • A small left-handed helix region (mostly occupied by glycine)
Ramachandran plot showing allowed combinations of phi and psi backbone dihedral angles. Dense clusters appear in the beta sheet region (upper left) and alpha helix region (middle left)
Ramachandran plot of allowed backbone angles. The upper-left cluster is the beta sheet region; the middle-left cluster is the alpha helix region. Steric clashes forbid most other combinations, which is why proteins repeatedly adopt the same secondary structures. Credit: Wikimedia Commons, CC BY 3.0 (D. Richardson via Dcrjsr)
Why is a peptide bond planar and rigid?
Click to reveal answer
The nitrogen lone pair delocalizes into the carbonyl C=O via resonance, giving the C-N bond about 40 percent double-bond character. Double bonds do not rotate. As a result, the six atoms of the peptide bond (Cα-C(=O)-N(H)-Cα) all lie in a single plane and cannot rotate around the C-N bond.
What are phi (Φ) and psi (Ψ) angles?
Click to reveal answer
They are the two rotatable backbone dihedral angles on either side of an alpha carbon. Phi is the rotation around the N-Cα bond; psi is the rotation around the Cα-C(=O) bond. The peptide bond itself (C-N) cannot rotate, so phi and psi define all the backbone flexibility of a polypeptide.
Why is trans preferred over cis for peptide bonds?
Click to reveal answer
Trans keeps the bulky R groups on opposite sides of the planar peptide bond, minimizing steric clashes. Cis puts them on the same side, which is energetically unfavorable. The exception is X-Pro bonds, where proline's cyclic side chain partially relieves the steric penalty and cis occurs about 5-10 percent of the time.
1.7

Primary Structure

The primary structure is the simplest thing in biology and the most important. It is just the order of amino acids in the chain, written from N-terminus to C-terminus. Nothing more.

Yet that one-dimensional string of letters determines everything: how the protein folds, what it binds, how fast it works, how long it lasts. Change one letter and you can destroy the whole protein - or, occasionally, build a new one.

How the Primary Structure is Specified

The gene dictates the primary sequence through the genetic code. Each codon (3 bases of mRNA) corresponds to one amino acid. The ribosome reads codons left to right and adds amino acids to the growing chain N-terminus first. Once translation finishes, the entire primary structure is set - nothing can change it without breaking a peptide bond.

Why Primary Structure is Hierarchically the Most Important

All other levels of structure are consequences of primary structure. The sequence “causes” the folding. Change the sequence and you change (potentially) everything downstream.

This is why biochemists say “the primary structure determines the tertiary structure.” Christian Anfinsen’s famous 1961 experiment showed this directly: ribonuclease was denatured (unfolded), then allowed to refold on its own in water. It regained its full enzymatic activity without any cellular help. The sequence carried all the information needed to fold.

Single Amino Acid Substitutions: Sickle Cell

The beta chain of adult hemoglobin is 146 amino acids long. In sickle cell disease, position 6 is changed from glutamate (E, negatively charged, hydrophilic) to valine (V, nonpolar, hydrophobic). One letter in 146. The rest of the sequence is identical.

Consequences of that single swap:

  • Valine on the surface creates a hydrophobic sticky patch.
  • At low oxygen, these hydrophobic patches on neighboring hemoglobin molecules stick together.
  • The hemoglobin polymerizes into long rigid fibers.
  • The fibers warp red blood cells into the characteristic sickle shape.
  • Sickled cells clog capillaries and cause ischemic pain, organ damage, and anemia.
Comparison of normal hemoglobin and sickle cell hemoglobin at the primary, secondary, and tertiary structure levels. A single amino acid substitution (glutamate to valine) at position 6 of the beta chain changes the tertiary structure and the shape of red blood cells
A single amino acid substitution in the hemoglobin beta chain (glutamate to valine at position 6) propagates through secondary, tertiary, and quaternary structure to distort the entire red blood cell. Primary structure controls everything downstream. Credit: OpenStax Biology 2e, CC BY 4.0
Scanning electron micrograph showing both normal disc-shaped red blood cells and sickled crescent-shaped red blood cells in a patient with sickle cell anemia
Real scanning electron micrograph of red blood cells. Normal cells are round and flexible; sickled cells are rigid crescents. All from one amino acid change in the hemoglobin primary sequence. Credit: Lumen Learning / OpenStax Biology 2e, CC BY 4.0

Determining Primary Structure in the Lab

Historically proteins were sequenced by Edman degradation: the N-terminal residue is cleaved off one at a time, identified by chromatography, and the process repeats. This is slow.

Modern methods use mass spectrometry. A protease (like trypsin) cuts the protein at specific residues (trypsin cuts after K or R), the fragments are analyzed by mass spec, and the sequence is reconstructed. A whole proteome can be identified this way.

What defines the primary structure of a protein?
Click to reveal answer
The linear sequence of amino acids from the N-terminus to the C-terminus, linked by covalent peptide bonds. The primary structure is encoded by the gene (via the genetic code) and contains all the information needed to determine how the protein will fold.
Why does the single Glu-to-Val mutation in hemoglobin beta cause sickle cell disease?
Click to reveal answer
Glutamate (E) is negatively charged and hydrophilic; valine (V) is nonpolar and hydrophobic. Placing a hydrophobic residue on the protein surface creates a sticky patch. At low oxygen, these patches on neighboring hemoglobins associate, polymerizing hemoglobin into rigid fibers that distort red blood cells into sickle shapes.
What does "primary structure determines tertiary structure" mean, and whose experiment demonstrated it?
Click to reveal answer
It means the amino acid sequence alone contains enough information to specify how a protein folds into its 3D shape. Christian Anfinsen showed this by denaturing ribonuclease and letting it refold in water without any cellular machinery - the enzyme regained full activity, proving the sequence alone encoded the fold.
1.8

Secondary Structure

After translation, the polypeptide chain begins to fold as it leaves the ribosome. The first folds are local - not the whole protein collapsing, just small regions snapping into repeating patterns. These local patterns are secondary structure, and they come in two flavors: alpha helices and beta pleated sheets.

Both are held together by the same force: hydrogen bonds between backbone amide groups. Not R groups. The backbone N-H of one residue hydrogen bonds to the backbone C=O of another residue a few positions away. R groups do not participate - they stick out to the side.

The Alpha Helix

The alpha helix is a right-handed corkscrew. The backbone spirals around an imaginary central axis. Every residue’s carbonyl oxygen (C=O) hydrogen bonds to the amide N-H exactly four residues ahead in the chain. Side chains project outward, away from the helix core.

Key numbers to know:

  • 3.6 residues per turn
  • 5.4 Å (0.54 nm) per turn - so 1.5 Å rise per residue
  • Hydrogen bond is from residue n to residue n+4
  • Most alpha helices are right-handed
Diagrams of the alpha helix (right-handed coil with hydrogen bonds running parallel to the helix axis) and the beta pleated sheet (zigzag strands linked by hydrogen bonds perpendicular to the strand direction), both with labeled hydrogen bonds and side chains
The alpha helix and beta pleated sheet - the two canonical secondary structures. Hydrogen bonds between backbone amide groups (not side chains) stabilize both. Credit: OpenStax Biology 2e, CC BY 4.0

Helix Breakers

Two amino acids disrupt alpha helices:

  • Proline has a ring that locks its Phi angle and removes the amide N-H needed for hydrogen bonding. Proline almost never appears inside a helix (only at the N-terminal “cap” position).
  • Glycine is too flexible. Its side chain is just hydrogen, so it samples too many conformations. It prefers turns and loops over the regular helix geometry.

The Beta Pleated Sheet

Beta sheets form between beta strands - segments of backbone running in an extended zigzag. Two or more strands line up next to each other and form hydrogen bonds sideways between adjacent strands. The result is a pleated sheet that looks like an accordion.

Beta sheets come in two flavors:

  • Parallel: both strands run in the same N-to-C direction. Hydrogen bonds are slightly skewed. Less stable than antiparallel.
  • Antiparallel: strands run in opposite N-to-C directions. Hydrogen bonds are linear (straight across). More stable.

Beta strands are separated in sequence but close in space. Between two adjacent strands in a sheet, there must be a connecting loop or “beta turn.”

Loops and Turns

Not every residue participates in a helix or sheet. In between, the backbone makes turns (short reversals of direction, often 3-4 residues) and loops (longer irregular segments). These regions often sit on the protein surface and contain active sites, binding loops, or antibody recognition loops.

Glycine and proline are over-represented in loops and turns. Glycine provides flexibility; proline enforces a sharp kink.

Why Secondary Structure Exists

Hydrogen bonding the backbone amides is energetically cheap and there are a lot of them in a polypeptide. If the protein did NOT form secondary structure, those H-bond donors (N-H) and acceptors (C=O) would need to find other partners - usually water. Burying the backbone in the protein interior requires first satisfying all those H-bonds internally. That is exactly what alpha helices and beta sheets do.

Between which atoms does the hydrogen bonding in secondary structure occur?
Click to reveal answer
Between the backbone carbonyl oxygen (C=O) of one residue and the backbone amide hydrogen (N-H) of another. The side chains (R groups) do not participate. In an alpha helix the bond is from residue n to residue n+4; in a beta sheet it runs between adjacent strands.
Which two amino acids are "helix breakers," and why?
Click to reveal answer
Proline (its cyclic side chain locks the phi angle and removes the N-H donor) and glycine (too flexible - its H side chain allows too many conformations to commit to the helix). Proline forces kinks; glycine is common in loops and turns.
What is the difference between a parallel and an antiparallel beta sheet, and which is more stable?
Click to reveal answer
In antiparallel beta sheets, neighboring strands run in opposite N-to-C directions, allowing linear hydrogen bonds that cross straight between strands - this is more stable. In parallel beta sheets, both strands run in the same direction, so the H-bonds tilt and are slightly less stable. Both are common in proteins.
1.9

Tertiary Structure

Primary structure is a flat sequence. Secondary structure is local folding (helices and sheets). Tertiary structure is the whole polypeptide chain folded into its final 3D shape, with the helices and sheets packed against each other and a stable core in the middle.

Tertiary structure is where the protein becomes functional. An enzyme’s active site is a tertiary feature. A binding pocket on a hormone receptor is a tertiary feature. If tertiary structure is wrong, the protein does not work.

The Five Forces That Stabilize Tertiary Structure

ForceWhat it isStrengthR groups involved
Hydrophobic interactionsNonpolar R groups cluster away from waterStrongest cumulative driver of foldingVal, Leu, Ile, Phe, Trp, Met, Ala
Disulfide bondsCovalent -S-S- between two cysteinesStrongest per-bond (covalent)Cys only
Salt bridges (ionic)Attraction between oppositely charged side chainsMediumAsp/Glu with Lys/Arg/His
Hydrogen bondsBetween polar side chains (or backbone)Medium-weakSer, Thr, Tyr, Asn, Gln, His
Van der Waals forcesWeak, cumulative close-range attractionsIndividually weak, collectively importantAll residues
Diagram illustrating the four main interactions that stabilize tertiary structure: hydrophobic clustering of nonpolar side chains, ionic bonds (salt bridges) between oppositely charged side chains, hydrogen bonds between polar side chains, and covalent disulfide bonds between cysteines
The main forces that stabilize tertiary structure. Hydrophobic clustering (at center) drives the overall fold, while disulfide bonds, ionic salt bridges, and hydrogen bonds lock the final shape in place. Credit: OpenStax Biology 2e, CC BY 4.0

Hydrophobic Interactions: the Main Engine

Nonpolar R groups cannot hydrogen bond with water. When a protein folds, exposed nonpolar residues disrupt the water’s hydrogen bond network, forcing water molecules to organize into “cages” around them (entropy penalty). Burying those nonpolar residues in a protein interior releases those ordered water molecules back into the bulk. That entropy gain is the driving force of folding.

Disulfide Bonds

Two cysteines can be oxidized to form a covalent -S-S- bridge. This is the strongest single bond stabilizing tertiary structure - a true covalent bond, not a weak interaction.

Disulfide bonds require an oxidizing environment. Inside the cytoplasm (reducing environment), they rarely form. In the endoplasmic reticulum and in extracellular space (oxidizing environments), they are common. That is why secreted proteins (insulin, antibodies, digestive enzymes) tend to be rich in disulfide bonds while cytosolic proteins are not.

Insulin structure showing two polypeptide chains (A chain and B chain) linked by two inter-chain disulfide bonds and a third intra-chain disulfide bond within the A chain
Insulin has two short polypeptide chains held together by three disulfide bonds (two between chains, one within the A chain). Disulfide bonds are common in secreted proteins. Credit: Lumen Learning / OpenStax Biology 2e, CC BY 4.0

Salt Bridges (Ionic Interactions)

Salt bridges form between a negatively charged side chain (Asp or Glu) and a positively charged one (Lys, Arg, His). They are weaker than covalent bonds but can have a significant effect on stability when buried inside the protein where no water can interfere.

Salt bridges are pH-sensitive. Lowering pH (protonating Asp/Glu) or raising pH (deprotonating Lys/Arg) removes the charge and breaks the bridge - one reason proteins denature at extreme pH.

Hydrogen Bonds and Van der Waals

Polar side chains (Ser, Thr, Tyr, Asn, Gln, His) can hydrogen bond with each other or with the backbone. Van der Waals forces (instantaneous dipoles) act between all atoms when they are close enough. Individually weak, but thousands of van der Waals contacts in a tightly packed protein interior add up.

What is the MOST important force driving protein folding, and why?
Click to reveal answer
The hydrophobic effect. Burying nonpolar side chains in the protein interior releases ordered water molecules that were trapped around exposed nonpolar surfaces. That increase in water's entropy drives the overall fold. Disulfide bonds are stronger per bond but are rare; hydrophobic interactions are everywhere in every protein and are the overall driver.
Why are disulfide bonds common in secreted proteins like insulin and antibodies but rare in cytoplasmic proteins?
Click to reveal answer
Disulfide bond formation requires an oxidizing environment. The cytoplasm is reducing (high glutathione), so -SH groups stay reduced. The endoplasmic reticulum and extracellular space are oxidizing, so cysteines can form -S-S- bridges. Secreted proteins traverse the ER and end up outside the cell, where disulfide bonds help stabilize them.
Between which amino acids would a salt bridge most likely form?
Click to reveal answer
Between a negatively charged side chain (aspartate or glutamate, D/E) and a positively charged one (lysine, arginine, or histidine; K/R/H). At physiological pH, these opposite charges attract. The interaction is called a salt bridge (ionic interaction) and it is pH-sensitive - protonation or deprotonation of either partner breaks it.
1.10

Quaternary Structure

Some proteins are lone wolves - one polypeptide chain folds, does its job, and that is the whole story (myoglobin, for example). Other proteins are teams - two or more polypeptide chains assemble into a functional complex. That team assembly is the quaternary structure.

Only proteins with more than one polypeptide chain have a quaternary structure. A single-chain protein has primary, secondary, and tertiary structure. Add a second chain and you have quaternary.

The Vocabulary

  • Monomer: single polypeptide chain.
  • Dimer: two subunits. A homodimer has two identical subunits; a heterodimer has two different ones.
  • Trimer: three subunits (e.g., collagen is a triple helix).
  • Tetramer: four subunits (e.g., hemoglobin).
  • Oligomer: general term for any multisubunit complex.

Hemoglobin: the Classic Example

Adult hemoglobin is a heterotetramer of two alpha chains and two beta chains (α2β2). Each subunit has its own tertiary structure and binds one heme group. The four subunits fit together in a compact roughly tetrahedral arrangement.

3D cartoon rendering of hemoglobin showing four globin subunits (two alpha in red-orange, two beta in blue) each cradling a heme group in the center of the tetramer
Hemoglobin, the classic quaternary-structure protein. Two alpha and two beta subunits (α2β2) assemble around four heme groups to create the oxygen carrier. Credit: Wikimedia Commons / Richard Wheeler (Zephyris), CC BY-SA 3.0 (from PDB 1GZX)
Progression from primary amino acid sequence to secondary alpha helix to tertiary folded single chain (beta globin polypeptide) to quaternary hemoglobin molecule with four subunits and heme groups
The four levels of protein structure, using hemoglobin as the example. Primary (sequence) builds secondary (helix), which folds into tertiary (one globin chain), which assembles into quaternary (full hemoglobin tetramer). Credit: OpenStax Biology 2e, CC BY 4.0

Why Quaternary Structure Matters

Two big reasons quaternary structure shows up all over biology:

  1. Cooperativity. When one subunit binds a ligand, it can change its shape and transmit that change to the others, making them more (or less) eager to bind. Hemoglobin does this with oxygen: binding O2 at one subunit pulls the other three into their high-affinity state. This is why the hemoglobin-oxygen curve is sigmoidal while single-subunit myoglobin’s curve is hyperbolic.

  2. Allosteric regulation. A molecule can bind at one site and affect activity at a distant site on another subunit. This is how many enzymes are controlled. Multisubunit proteins have more surface area and more interfaces for regulatory molecules to latch onto.

Subunit Interfaces Use the Same Forces as Tertiary Structure

Quaternary assembly uses the same bond types that build tertiary structure: hydrophobic contacts at the subunit interface, ionic interactions, hydrogen bonds, and sometimes disulfide bonds that bridge two chains (like the inter-chain disulfides in insulin and immunoglobulins). There is no new chemistry at the quaternary level - just a bigger playing field.

Simple schematic of an IgG antibody showing two heavy chains and two light chains arranged in a Y-shape held together by interchain disulfide bonds
IgG antibody schematic. Two heavy chains and two light chains assemble into a Y-shape held together by disulfide bonds - a textbook example of quaternary structure. Credit: Servier Medical Art, CC BY 4.0
What defines quaternary structure, and which proteins have it?
Click to reveal answer
Quaternary structure is the assembly of two or more polypeptide chains (subunits) into a functional complex. Only multisubunit proteins have quaternary structure. Examples: hemoglobin (α2β2 tetramer), antibodies (2 heavy + 2 light chains), collagen (triple helix). Single-chain proteins like myoglobin do not have quaternary structure.
Why does hemoglobin show a sigmoidal O2 binding curve while myoglobin shows a hyperbolic one?
Click to reveal answer
Hemoglobin has four cooperating subunits; myoglobin has only one. When one hemoglobin subunit binds O2, it nudges the other subunits into a higher-affinity state (positive cooperativity), giving the S-shaped sigmoidal curve. Myoglobin has no subunits to cooperate with, so its binding follows a simple hyperbolic curve.
A protein runs as one band on native gel but as two bands of different sizes on SDS-PAGE with a reducing agent. What does this tell you?
Click to reveal answer
The protein has quaternary structure with two different subunits held together by disulfide bonds. Native gel keeps the complex intact (one band). SDS-PAGE with a reducing agent (like DTT or β-mercaptoethanol) denatures the protein and breaks disulfide bonds, separating the complex into its individual subunits. The fact that two bands appear shows two non-identical chains.
1.11

Folding and Denaturation

A polypeptide emerges from the ribosome as a loose chain. Over seconds to minutes, it folds into a specific 3D shape called the native state. That folded structure is what does the job - whether the job is catalyzing a reaction, carrying oxygen, or supporting a cell.

The four levels, and the bond that holds each one

Protein structure
One polypeptide, four stages. Each stage is the one before it, folded further. 1° Primary N C side chains, not yet doing anything amino acids in a fixed order 2° Secondary α-helix H bonds run along the axis, i to i+4 β-sheet · arrows point N → C antiparallel: H bonds run sideways 3° Tertiary S–S bridge · amber = buried core 4° Quaternary α β β α hemoglobin · α₂β₂ · red disc = heme Primary the sequence Peptide bonds
Amino acids in a fixed order, N-terminus to C-terminus. Written by the gene, and nothing later can change it.
Secondary local shapes Backbone hydrogen bonds
α-helix and β-pleated sheet. Each backbone N–H bonds to a C=O four residues along, so the side chains take no part at all.
Tertiary the whole fold Side chains, plus disulfides
Helices and sheets packed into one domain, held by a hydrophobic core, hydrogen bonds, ionic pairs, and cysteine disulfide bridges.
Quaternary subunits together Same interactions, new chains
Two or more folded subunits assembling. Only some proteins have it: hemoglobin does, myoglobin does not.
Denaturation: what actually breaks Primary structure survives all of it Heat · extreme pH
Breaks hydrogen bonds and ionic pairs, so quaternary, tertiary, and secondary all go.
Urea · detergent
Breaks the hydrophobic core by making the solvent friendlier to nonpolar side chains.
β-mercaptoethanol
The only one that touches covalent bonds: it reduces disulfide bridges back to free thiols.
1

Scroll sideways to see the whole map.

α-helix: a coiled ribbon β-strand: an arrow, N to C Covalent: peptide bonds and disulfides Buried nonpolar side chains
One molecule, four descriptions of it. Each level is the previous one folded further, held by interactions the sequence had already decided. That is why a single substituted residue, as in sickle cell, can change the shape of the whole assembled protein.

Break the native shape and the protein stops working. Breaking the shape is called denaturation. Sometimes denatured proteins can refold. Often they cannot.

The Native State

The native state is the conformation where the protein has its lowest free energy under physiological conditions. Usually this corresponds to the functional fold. The native state is:

  • Stable enough to survive normal cellular conditions
  • Often flexible enough to allow catalysis or conformational changes
  • Not always the most thermodynamically stable form possible - some proteins fold into metastable, kinetically trapped states, especially membrane proteins and amyloids

Folding happens in milliseconds to seconds for most proteins, which is remarkable given how many possible conformations a 150-residue chain could explore.

The Thermodynamics of Folding

Folding looks like order emerging from disorder, which seems to violate the second law of thermodynamics. It does not - you just have to count the water.

The governing equation is ΔG = ΔH - TΔS. For folding to be spontaneous, ΔG must be negative. Here is what each term actually contributes:

This is the hydrophobic effect, and it is the #1 driver of folding. It explains why:

  • Hydrophobic residues end up buried in the protein core
  • Polar and charged residues stay on the surface, where they contact water
  • Heating can denature a protein (heat raises TΔS for the unfolded state even more, tipping the balance)

Chaperones

Inside cells, many proteins cannot fold on their own - the cytoplasm is too crowded and some intermediates have exposed hydrophobic patches that would stick to neighboring proteins and aggregate. Chaperone proteins bind to newly synthesized polypeptides and keep them isolated while they fold.

The main chaperone families (no need to memorize the exact names):

  • Hsp70 family - binds nascent chains, prevents aggregation
  • Hsp60 / GroEL-GroES family - provides an isolated barrel-shaped chamber where a single polypeptide can fold

What Denatures a Protein

Denaturation is the loss of the folded 3D structure. The primary sequence (covalent peptide bonds) is almost always preserved - only the weak forces that stabilize folding get disrupted.

| Agent | Mechanism |
|-------|-----------|
| Heat | Vibrations overcome hydrogen bonds and hydrophobic packing |
| Low pH (acid) | Protonates carboxyls, disrupts salt bridges, protonates histidines |
| High pH (base) | Deprotonates amines and tyrosines, disrupts salt bridges |
| Urea and guanidinium chloride | Chaotropes - compete for H-bonds with water, destabilizing the hydrophobic effect |
| Detergents (SDS) | Coat the protein surface, break hydrophobic contacts |
| Reducing agents (β-ME, DTT) | Break disulfide bonds |
| Heavy metals | Bind to cysteines, ionic groups, disrupt folding |
| Mechanical agitation | Exposes hydrophobic interior (whipping egg whites) |

Denaturation: Reversible or Irreversible?

Sometimes denatured proteins spontaneously refold when the denaturant is removed. Ribonuclease is the classic case - Anfinsen showed that ribonuclease denatured by urea and DTT refolded to full activity when the chemicals were washed out. The sequence held all the information.

More often, denaturation is effectively irreversible:

  • Egg whites (fried or whipped) cannot be un-denatured
  • Boiled milk cannot refold its proteins
  • Most denatured proteins aggregate before they get a chance to refold, because exposed hydrophobic patches stick together
Does denaturation break primary structure?
Click to reveal answer
No. Denaturation destroys secondary, tertiary, and quaternary structure by disrupting weak forces (hydrogen bonds, hydrophobic contacts, ionic bonds). The covalent peptide bonds of the primary sequence remain intact. To break primary structure you need hydrolysis of peptide bonds - typically by strong acid, strong base, or a protease.
What do chaperone proteins actually do, and do they determine the final fold?
Click to reveal answer
Chaperones prevent premature aggregation and provide an environment where a polypeptide can fold without interference from other proteins. They do NOT dictate the final conformation - the amino acid sequence contains all the information needed for folding (Anfinsen’s dogma). Chaperones simply protect the folding process from the crowded cellular environment.
How does SDS denature proteins in SDS-PAGE?
Click to reveal answer
SDS is an amphipathic detergent. Its long hydrophobic tail coats the protein’s nonpolar interior, breaking the hydrophobic interactions that hold tertiary structure together. Its negatively charged sulfate head gives all proteins the same charge-to-mass ratio, so they separate by size alone on the gel. Reducing agents are usually added alongside SDS to also break disulfide bonds.
1.12

Conjugated Proteins

Many proteins cannot do their job with just 20 amino acids. They recruit a non-protein helper - a metal ion, a vitamin derivative, a sugar chain, a heme group. A protein that carries a non-protein component is called a conjugated protein.

The protein part is the apoprotein (without the helper). Together with the helper, it is the holoprotein (the complete, functional form). If the apoprotein cannot function alone, the helper is essential - and the helper itself gets a name depending on how tightly it binds.

Prosthetic Groups vs. Cofactors vs. Coenzymes

Three related terms that get confused:

  • Cofactor: any non-protein helper. Includes both metal ions and organic molecules.
  • Coenzyme: an organic cofactor. Most are vitamin-derived (NAD+, FAD, CoA, biotin).
  • Prosthetic group: a cofactor that is permanently and tightly bound to the protein (often covalently). Heme is a prosthetic group.

The simple way to keep this straight: prosthetic group = glued on; cofactor = the general category; coenzyme = an organic cofactor that may bind loosely.

Major Types of Conjugated Proteins

TypeAttached groupExample proteinFunction of attachment
HemoproteinHeme (iron porphyrin)Hemoglobin, myoglobin, cytochromesBinds O2 or transfers electrons
GlycoproteinCarbohydrate chainsAntibodies, cell surface receptors, many secreted proteinsCell recognition, stability
LipoproteinLipid cargoHDL, LDL, VLDLTransports lipids in blood
MetalloproteinMetal ion (Fe, Zn, Cu, Mg)Carbonic anhydrase (Zn), superoxide dismutase (Cu, Zn)Catalysis, structure
PhosphoproteinPhosphate groupCasein (milk), many regulatory proteinsRegulation, mineral binding
NucleoproteinNucleic acidRibosome, chromatin (histones + DNA)Genetic information handling
Simple cartoon illustration of a folded protein, representing the apoprotein portion of a conjugated protein before a prosthetic group is added
A folded polypeptide ready to accept its prosthetic group. The protein scaffold (apoprotein) cradles the non-protein helper, whether that is heme, a metal ion, a carbohydrate, or a lipid. Credit: Servier Medical Art, CC BY 4.0

Hemoproteins

Heme is an iron ion held inside a porphyrin ring. It is a prosthetic group in:

  • Hemoglobin and myoglobin - reversibly bind O2
  • Cytochromes - shuttle electrons in the electron transport chain
  • Catalase - breaks down hydrogen peroxide

All heme proteins use the iron’s ability to switch between Fe2+ and Fe3+ (cytochromes) or to hold Fe2+ in a special “ready” state for O2 binding (hemoglobin/myoglobin).

Glycoproteins

A carbohydrate chain (often branched) is attached to a serine, threonine, or asparagine side chain. Glycoproteins are everywhere:

  • Cell-surface receptors and adhesion molecules
  • Antibodies (IgG is about 4 percent carbohydrate)
  • Blood-type antigens (ABO antigens are glycolipid/glycoprotein differences)
  • Most secreted proteins have sugar decorations

Lipoproteins

Proteins that wrap around a core of cholesterol and triglycerides to transport lipids through the aqueous blood. The major classes (HDL, LDL, IDL, VLDL, chylomicrons) differ in lipid-to-protein ratio, which affects their density. More lipid = less dense. You will see these in detail in the lipid chapter.

Metalloproteins

A metal ion is held in the active site by side chain ligands (usually histidine, cysteine, aspartate, or glutamate). Metal ions catalyze reactions that the 20 amino acids alone cannot - redox chemistry (Fe, Cu), Lewis-acid catalysis (Zn), and structural stabilization (Mg, Zn fingers in transcription factors).

What is the difference between an apoprotein and a holoprotein?
Click to reveal answer
The apoprotein is the protein alone, without its cofactor or prosthetic group - often inactive. The holoprotein is the complete functional form, with the cofactor bound. Example: apohemoglobin (no heme) is non-functional; hemoglobin (with four hemes) is the holoprotein that carries oxygen.
What is the difference between a prosthetic group and a coenzyme?
Click to reveal answer
Prosthetic groups are permanently (often covalently) bound cofactors that stay with the protein (e.g., heme in hemoglobin, FAD in succinate dehydrogenase). Coenzymes are usually organic (often vitamin-derived) cofactors that bind loosely - they can come and go with each catalytic cycle (e.g., NAD+ in most dehydrogenases). Both are cofactors.
Why might a recombinant human protein made in E. coli be less active than the same protein purified from human cells?
Click to reveal answer
Bacteria do not perform the same post-translational modifications as eukaryotes. They especially cannot add eukaryotic-style glycosylation. If the human protein is a glycoprotein that needs specific sugar chains for folding, stability, or activity, the bacterial version will be missing them and may misfold or be cleared quickly in vivo.