A nucleotide has three parts: a nitrogenous base, a five-carbon sugar, and a phosphate (or phosphates). Add the phosphate and it becomes a nucleoTIDE. Remove the phosphate and it is just a nucleoSIDE (sugar + base).
A nucleotide: base + sugar + phosphate. In a polynucleotide chain, the phosphate links the 5' carbon of one sugar to the 3' carbon of the next. Credit: Wikimedia Commons, CC BY-SA
The Five Bases
Two families of bases, distinguished by their ring structure:
Purines: two fused rings. Adenine (A) and guanine (G).
Pyrimidines: one ring. Cytosine (C), thymine (T), and uracil (U).
Thymine is in DNA only. Uracil is in RNA only. A, G, and C are in both.
Sugars: DNA vs. RNA
DNA uses 2-deoxyribose - C2 has no hydroxyl (just -H).
RNA uses ribose - C2 has a hydroxyl (-OH).
That single hydroxyl difference makes RNA far less chemically stable than DNA. The 2’-OH of RNA can attack the phosphodiester backbone intramolecularly, cleaving the chain. DNA, lacking this attacker, is much more stable - appropriate for long-term information storage.
Nucleoside vs. Nucleotide Naming
Base
Nucleoside (base + sugar)
Nucleotide (add phosphate)
Adenine
Adenosine (RNA) / Deoxyadenosine (DNA)
AMP, ADP, ATP
Guanine
Guanosine / Deoxyguanosine
GMP, GDP, GTP
Cytosine
Cytidine / Deoxycytidine
CMP, CDP, CTP
Thymine
Thymidine
TMP, TDP, TTP
Uracil
Uridine
UMP, UDP, UTP
The Phosphodiester Backbone
Nucleotides polymerize through phosphodiester bonds: the 5’-phosphate of one nucleotide attaches to the 3’-OH of the next. This gives the backbone its directionality. Every DNA or RNA strand has a 5’ end (free phosphate) and a 3’ end (free -OH).
Polymerases always add new nucleotides to the 3’-OH, so synthesis runs 5’ → 3’. Reading and writing conventions follow this: a sequence written as ATCG means the 5’ end has A and the 3’ end has G.
What three components make up a nucleotide, and which are present in a nucleoside?
Click to reveal answer
A nucleotide has three parts: a nitrogenous base, a five-carbon sugar, and one or more phosphate groups. A nucleoside has just the base and sugar (no phosphate). Adding a phosphate converts a nucleoside to a nucleotide.
Which bases are purines and which are pyrimidines?
Click to reveal answer
Purines (two-ring bases): adenine (A) and guanine (G). Pyrimidines (one-ring bases): cytosine (C), thymine (T, DNA only), and uracil (U, RNA only). Mnemonic: "Pure As Gold" for purines A and G.
In which direction do DNA and RNA polymerases synthesize new strands, and why?
Click to reveal answer
5' to 3'. The 3'-OH of the growing strand attacks the alpha-phosphate of an incoming 5'-triphosphate nucleotide, forming a new phosphodiester bond and releasing pyrophosphate. Without a free 3'-OH, polymerization cannot proceed - which is the basis of Sanger sequencing (adding 2',3'-dideoxy nucleotides terminates the chain).
DNA is a right-handed double helix of two antiparallel polynucleotide strands. The backbones run on the outside; the bases pair in the center.
What a nucleotide is, and what holds two strands together
DNA structure
1
Scroll sideways to see the whole map.
Purines: two rings (A and G) Pyrimidines: one ring (C, T, U) Three hydrogen bonds: G-C Sugar-phosphate backbone
Antiparallel, and what followsOne strand runs 5' to 3' and its partner runs 3' to 5'. Because every polymerase can only build toward a 3' end, that single geometric fact forces the leading and lagging strands in replication, sets the direction of transcription, and defines what a template strand is.
Reading a sequenceBy convention a sequence is written 5' to 3' unless stated otherwise. The complement of 5'-ATGC-3' is 3'-TACG-5', which written conventionally is 5'-GCAT-3'. Half the mistakes on this topic are from forgetting to flip.
Purines and pyrimidinesPURe As Gold: purines are Adenine and Guanine, and they have two rings. CUT the PY: pyrimidines are Cytosine, Uracil, and Thymine, with one ring. Every pair is one of each, which is what keeps the helix a constant width.
Two facts carry the whole chapter: the strands are antiparallel, and G-C is held by three bonds where A-T is held by two. The first dictates how DNA is copied and read; the second dictates how firmly it is held together.The right-handed DNA double helix. Sugar-phosphate backbones are on the outside; base pairs are on the inside. Strands run antiparallel - one 5’ to 3’, the other 3’ to 5’. Credit: Wikimedia Commons, CC BY-SA
Base Pairing
The two strands are held together by hydrogen bonds between paired bases. Watson-Crick base pairing follows strict rules:
A pairs with T via 2 hydrogen bonds.
G pairs with C via 3 hydrogen bonds.
A purine always pairs with a pyrimidine, keeping the helix diameter constant.
Antiparallel Strands
The two strands run in opposite directions. One strand is 5’ → 3’ top to bottom; the other is 5’ → 3’ bottom to top. Each base on one strand pairs with the complementary base directly across from it on the other strand.
This antiparallel arrangement has a major consequence: since DNA polymerases only synthesize 5’ → 3’, the two strands at a replication fork must be copied differently. One is made continuously (leading strand), one is made in fragments (lagging strand).
Grooves
The helix has two grooves where proteins can read the base sequence without unwinding the helix:
Major groove: wider; exposes more of the base edges. Most transcription factors recognize DNA sequences here.
Minor groove: narrower; less information. Some specialized proteins read sequences here.
Chargaff’s Rules
Before Watson and Crick, Erwin Chargaff showed that for any double-stranded DNA sample:
The amount of A equals the amount of T, and the amount of G equals the amount of C.
The percent of purines (A + G) equals the percent of pyrimidines (T + C), which is always 50%.
These relationships make sense only if A pairs with T and G pairs with C. Chargaff’s data was a critical clue in Watson and Crick’s 1953 model.
DNA Denaturation, Reannealing, and Hybridization
The hydrogen bonds between paired bases are much weaker than the covalent backbone, so you can pull the two strands apart without breaking the sequence. This is denaturation (sometimes called “melting”). Heat is the usual cause in the lab; extreme pH or chaotropes (urea, formamide) also work.
Melting temperature (Tm): the temperature at which half of the DNA is single-stranded and half is double-stranded. Higher GC content = higher Tm (three H-bonds per GC vs. two per AT). Longer DNA and higher salt also raise Tm.
Reannealing: if cooling is slow after denaturation, the separated strands will find their complements again and reform the double helix. The optimal reannealing temperature is about 20-25°C below Tm.
Hybridization: single-stranded DNA (or RNA) from one source pairs with complementary single-stranded nucleic acid from another source. This is how probes work in Southern and Northern blots, how primers find their targets in PCR, and how microarrays read gene expression.
Chromatin Structure in the Nucleus
A single human cell contains about 2 meters of DNA packaged into a nucleus only a few microns across. The packaging solution is chromatin - DNA wound around proteins and folded hierarchically.
Histones: small basic proteins rich in lysine and arginine (positively charged to bind the negative DNA backbone). Five histone types: H2A, H2B, H3, H4, and the linker H1.
Nucleosome: ~147 bp of DNA wrapped about 1.65 turns around a histone octamer (two copies each of H2A, H2B, H3, H4). Nucleosomes are the fundamental repeating unit of chromatin, often described as “beads on a string.”
30-nm fiber: nucleosomes coil into thicker fibers, with H1 stabilizing the coil.
Higher-order folding produces the condensed metaphase chromosome, which is ~10,000-fold shorter than the naked DNA.
Euchromatin vs. Heterochromatin
Euchromatin: loosely packed, transcriptionally active. Appears lighter under microscopy.
Heterochromatin: densely packed, transcriptionally silent. Appears darker. Constitutive heterochromatin (always condensed, like centromeres and telomeres) vs. facultative heterochromatin (tissue- or time-specific silencing, like the inactivated X chromosome).
Telomeres and Centromeres
Telomeres: repetitive sequences (TTAGGG in humans) at chromosome ends. They cap the ends to prevent fusion and degradation, and they shorten with each round of replication unless telomerase extends them. Shortening is linked to cellular aging; telomerase reactivation is a hallmark of many cancers.
Centromeres: heterochromatic repetitive regions in the middle of each chromosome that anchor sister chromatids together and serve as the assembly site for the kinetochore during mitosis.
Single-Copy vs. Repetitive DNA
Only about 1-2 percent of the human genome codes for protein. The rest includes regulatory elements, introns, and large stretches of repetitive DNA.
Single-copy DNA: most protein-coding genes. Present as one (or a few) copies per genome.
Repetitive DNA: tandem repeats (like centromeric and telomeric repeats) and interspersed repeats (like transposon-derived SINEs and LINEs). Some repetitive regions are structural; others are evolutionary relics.
How many hydrogen bonds hold each Watson-Crick base pair?
Click to reveal answer
A-T pairs are held by 2 hydrogen bonds. G-C pairs are held by 3 hydrogen bonds. This is why GC-rich DNA has a higher melting temperature than AT-rich DNA - it takes more energy to break the stronger GC hydrogen bonding.
A double-stranded DNA sample is 30% adenine. What percentage is guanine?
Click to reveal answer
20%. By Chargaff’s rules, %A = %T, so T = 30%. A + T together = 60%. The remaining 40% is G + C, split equally as %G = %C, so G = 20% and C = 20%.
Why must the two strands of a DNA double helix be antiparallel?
Click to reveal answer
Base pairing geometry requires it. The hydrogen-bond donors and acceptors on A, T, G, and C only line up properly when one strand runs 5’ → 3’ and the other runs 3’ → 5’. This antiparallel geometry forces the asymmetric leading/lagging strand replication pattern.
What is a nucleosome, and what holds it together?
Click to reveal answer
A nucleosome is ~147 bp of DNA wrapped about 1.65 turns around a histone octamer (two copies each of H2A, H2B, H3, H4). The attraction is electrostatic: histones are rich in positively charged lysine and arginine residues, which bind the negatively charged phosphate backbone. H1 linker histones stabilize the compaction between nucleosomes.
What is the melting temperature (Tm) of DNA, and what raises it?
Click to reveal answer
Tm is the temperature at which half of the DNA is single-stranded and half is double-stranded. Raising Tm: higher GC content (3 H-bonds per GC vs. 2 per AT), longer DNA, higher salt (shields phosphate repulsion). Controlling Tm is fundamental to PCR primer design, hybridization probes, and blotting.
What is the difference between euchromatin and heterochromatin?
Click to reveal answer
Euchromatin is loosely packed and transcriptionally active - the “working” portion of the genome. Heterochromatin is densely packed and transcriptionally silent, either constitutively (telomeres, centromeres) or facultatively (like the inactivated X). Chromatin state is regulated by DNA methylation and histone modifications.
DNA replication produces two identical double helices from one original. The key experiment (Meselson-Stahl, 1958) showed that replication is semiconservative - each daughter molecule contains one original strand and one newly synthesized strand.
The replication fork, enzyme by enzyme
DNA replication
1
Scroll sideways to see the whole map.
Leading strand: continuous Lagging strand: in fragments RNA primer Parent strands
Why there is a lagging strand at allDNA polymerase can only add to a 3' end, so it can only build 5' to 3'. The two template strands run in opposite directions, so as the fork opens, one template presents its 3' end continuously and the other keeps presenting new 5' ends. The second one has to be copied in short backwards pieces.
Two different exonucleasesDNA polymerase III proofreads with 3'→5' exonuclease, removing a base it just added wrongly. DNA polymerase I uses 5'→3' exonuclease to chew forwards through the RNA primers and replace them. Same word, opposite directions, different jobs.
The end-replication problemThe very last primer on the lagging strand cannot be replaced, because there is no upstream 3' end to extend from. Chromosomes therefore shorten every division, which is what telomeres buffer and what telomerase reverses in stem cells and in most cancers.
Semi-conservative, bidirectional, and always 5' to 3'. Those three constraints between them force everything else on this map: the primers, the fragments, the two polymerases, the ligase, and the fact that chromosome ends are a problem at all.
The Replication Fork
Replication begins at an origin of replication. The double helix is unwound, creating a Y-shaped replication fork. DNA polymerase synthesizes new DNA on both template strands - but each strand is made differently because of the antiparallel geometry.
Leading strand: synthesized continuously 5’ to 3’ toward the fork.
Lagging strand: synthesized discontinuously in short 5’ to 3’ fragments called Okazaki fragments. These are later joined together.
Key Enzymes
| Enzyme | Job |
|--------|-----|
| Helicase | Unwinds the double helix at the fork |
| Topoisomerase (gyrase in bacteria) | Relieves supercoiling ahead of the fork |
| Single-strand binding proteins (SSBs) | Prevent reannealing of the separated strands |
| Primase | Synthesizes short RNA primers (~10 nt) that DNA polymerase can extend |
| DNA polymerase III (bacteria) / δ, ε (eukaryotes) | Main replicative polymerase. Adds dNTPs to the 3’-OH of the growing strand |
| DNA polymerase I (bacteria) | Removes RNA primers, fills in gaps with DNA |
| DNA ligase | Seals nicks between Okazaki fragments |
Fidelity
DNA polymerases have very low error rates - about one mistake per 107 bases. Several mechanisms contribute:
Base selection: the active site is shaped to fit correct base pairs much better than mismatches.
Proofreading: most replicative polymerases have 3’ to 5’ exonuclease activity that removes a mispaired base just after it is added, then tries again.
Mismatch repair: post-replication, a separate system scans newly made DNA for any surviving mismatches and corrects them (see DNA repair section).
Together these systems bring the overall error rate down to about one mistake per 109 bases.
What is semiconservative DNA replication, and how was it shown experimentally?
Click to reveal answer
Each daughter DNA molecule consists of one original (“parental”) strand and one newly synthesized (“daughter”) strand. Meselson and Stahl grew E. coli in 15N (heavy nitrogen), switched to 14N (light nitrogen), and tracked the DNA density. After one round of replication, all DNA had intermediate density (one heavy strand + one light strand), consistent only with semiconservative replication.
Why must the lagging strand be synthesized discontinuously?
Click to reveal answer
DNA polymerase only adds nucleotides to a free 3’-OH, meaning it synthesizes 5’ to 3’. At the replication fork, one template is oriented so synthesis runs into the fork (continuous leading strand). The other template is oriented so 5’ to 3’ synthesis would run AWAY from the fork. The lagging strand must therefore be synthesized in short fragments (Okazaki fragments), each started with a new primer as more template is exposed. Ligase later joins them.
Why are RNA primers needed at the start of each new DNA fragment?
Click to reveal answer
DNA polymerases cannot start a new strand from scratch - they can only extend an existing 3’-OH. Primase, a specialized RNA polymerase, does not need a primer and lays down a short RNA primer on the template. DNA polymerase then extends that primer. After replication, the RNA primers are removed and replaced with DNA (by polymerase I in bacteria, or by a separate enzyme in eukaryotes).
DNA gets damaged constantly. UV light, chemical mutagens, reactive oxygen species, and spontaneous hydrolysis all cause lesions. Without repair, mutations would accumulate rapidly. Several distinct repair pathways exist, each specialized for a different kind of damage.
Proofreading
Part of replication itself. Most DNA polymerases have a 3’ to 5’ exonuclease activity that reads back after placing a new nucleotide. If the new base is not correctly paired, the polymerase chops it off and tries again. This catches most errors during synthesis.
Mismatch Repair (MMR)
Fixes the rare errors that escape proofreading. The mismatch repair system scans freshly replicated DNA, recognizes mispaired bases, and replaces the incorrect nucleotide on the newly synthesized strand (not the template). How does it know which strand is new? In bacteria, the parental strand is methylated at certain sequences and the new strand is not yet methylated. In eukaryotes, the mechanism is less clear but probably uses nicks in the new strand.
MMR defects cause hereditary non-polyposis colorectal cancer (HNPCC, Lynch syndrome) - mutations in MLH1, MSH2, and other MMR genes.
Base Excision Repair (BER)
Repairs single damaged bases (like oxidized or deaminated ones). A DNA glycosylase removes the damaged base, leaving an abasic site. An endonuclease nicks the backbone. A polymerase fills the gap with a correct nucleotide, and ligase seals the nick.
Nucleotide Excision Repair (NER)
Repairs bulky lesions that distort the helix, like UV-induced thymine dimers. Endonucleases cut out a short segment (~12-24 nucleotides) around the damage. Polymerase fills the gap, ligase seals it.
NER defects cause xeroderma pigmentosum (XP) - patients cannot repair UV damage, develop skin cancers as children, and must avoid sunlight entirely.
Double-Strand Break Repair
Double-strand breaks are the worst kind of damage - both strands are cut. Two repair paths:
Homologous recombination (HR): uses an identical sister chromatid as a template. Accurate but requires an intact copy, so it is restricted to the S and G2 phases of the cell cycle. BRCA1 and BRCA2 proteins are essential components. Mutations in BRCA21 cause hereditary breast and ovarian cancer.
Non-homologous end joining (NHEJ): directly glues the two ends back together without a template. Fast but error-prone - small insertions or deletions often result. Used throughout the cell cycle.
Clinical Examples
Disease
Defective pathway
Consequence
Xeroderma pigmentosum
Nucleotide excision repair
UV sensitivity, early skin cancer
Lynch syndrome (HNPCC)
Mismatch repair
Colorectal, endometrial cancer
BRCA21 mutations
Homologous recombination
Hereditary breast/ovarian cancer
Ataxia-telangiectasia
DNA damage sensing (ATM)
Neurodegeneration, immunodeficiency, cancer
What kind of DNA damage does nucleotide excision repair handle, and what disease results from defective NER?
Click to reveal answer
NER repairs bulky helix-distorting lesions, most notably UV-induced pyrimidine dimers (like thymine-thymine dimers). Xeroderma pigmentosum (XP) results from defective NER: patients cannot repair UV damage, develop multiple skin cancers in childhood, and must avoid sunlight.
What is the difference between homologous recombination and non-homologous end joining?
Click to reveal answer
Both repair double-strand breaks. Homologous recombination uses an intact sister chromatid as a template, producing error-free repair but requiring an identical DNA copy (S/G2 only). Non-homologous end joining directly ligates the two broken ends without a template - fast but error-prone, often resulting in small indels. NHEJ is used throughout the cell cycle.
How does mismatch repair know which of the two strands to correct?
Click to reveal answer
MMR must avoid editing the correct (parental) strand. In bacteria, the parental strand carries methyl groups on adenines at GATC sequences; the newly synthesized strand is not yet methylated. MMR acts on the unmethylated strand. In eukaryotes, the mechanism is different and probably involves nicks or other markers in the new strand, but the principle is the same: the system identifies which strand is new and preserves the template.
Restriction enzymes are bacterial proteins that cut DNA at specific sequences. Bacteria evolved them to slice up invading viral DNA. Molecular biologists hijack them to cut and paste DNA in the lab - the foundation of modern genetic engineering.
Which technique answers which question
Biotechnology
1
Scroll sideways to see the whole map.
Make more of it Find or read it Move it somewhere Change it
Why restriction sites are palindromesA restriction site such as GAATTC reads the same 5' to 3' on both strands. That is what lets one enzyme, working as a dimer, cut both strands at equivalent positions and leave matching single-stranded overhangs on each side.
Why PCR needed a hot-spring bacteriumEach cycle starts by heating to 95° to separate the strands, which would destroy an ordinary polymerase. Taq, from a thermophile, survives it, and that single fact is what made the technique practical.
cDNA and why it mattersReverse transcriptase turns mRNA into cDNA, which has no introns because the message was already spliced. That is how a human gene is expressed in bacteria, which cannot splice: you give them the edited version.
Match the verb in the question to the technique. Amplify means PCR. Separate means a gel. Detect means a blot. Read means sequencing. Insert means cloning. Edit means CRISPR. Almost every question on this topic is that translation.
Recognition Sequences
Restriction enzymes recognize and cut palindromic sequences. A palindrome in DNA reads the same 5’ to 3’ on both strands. For example, EcoRI recognizes:
5'-G A A T T C-3'3'-C T T A A G-5'
Read the top strand 5’ to 3’ (GAATTC), then read the bottom strand 5’ to 3’ (from right to left on the page, GAATTC). Same sequence. That is a palindrome.
Sticky Ends and Blunt Ends
EcoRI cuts between G and A on each strand, leaving overhangs:
5'-G A A T T C-3'3'-C T T A A G-5'
These single-stranded overhangs (sticky ends) can base-pair with any other EcoRI-cut DNA. This is what makes cloning work - you cut a gene of interest with EcoRI, cut your plasmid vector with EcoRI, and the sticky ends recombine. Ligase seals the nicks.
Some enzymes (like EcoRV, SmaI) cut straight across, producing blunt ends. Blunt ends are less selective in ligation but still usable.
Plasmid Vectors
A plasmid is a small circular DNA molecule found naturally in bacteria. Cloning vectors are engineered plasmids with:
An origin of replication (ori) so the plasmid replicates in bacterial hosts.
A multiple cloning site (MCS) - a stretch of unique restriction sites where you can insert your gene.
A selectable marker (usually an antibiotic resistance gene) so bacteria that contain the plasmid can be selected.
To clone a gene:
Cut both the gene and the plasmid with the same restriction enzyme.
Mix fragments with DNA ligase.
Transform bacteria with the ligated product (e.g., by heat shock or electroporation).
Grow bacteria on a selective medium (with antibiotic) - only transformed bacteria with the plasmid survive.
Confirmed bacterial colonies now carry millions of copies of your insert.
Applications
Recombinant DNA technology enables:
Protein production: human insulin, growth hormone, and many other therapeutics are made by bacteria engineered to express the human gene.
Gene analysis: any gene can be cloned and studied - its sequence, expression, and function.
Genetic engineering: cloning plus modern tools (CRISPR, transgenic organisms) has built modern biotechnology.
What is a palindromic DNA sequence, and why do restriction enzymes recognize them?
Click to reveal answer
A palindrome reads the same 5’ to 3’ on both strands (e.g., GAATTC). Restriction enzymes are typically homodimers that bind as two copies arranged in opposite orientations. The palindromic sequence allows the two subunits to make identical contacts on each strand, enabling specific, symmetric cleavage.
Why are sticky ends preferred over blunt ends for cloning?
Click to reveal answer
Sticky ends have short single-stranded overhangs with known sequences that complementary-base-pair with any other DNA cut by the same enzyme. This positions the two fragments correctly before ligase seals the nicks, giving efficient and directional ligation. Blunt ends ligate less efficiently and in either orientation.
What three essential features does a cloning plasmid contain?
Click to reveal answer
(1) An origin of replication so the plasmid replicates in the bacterial host. (2) A multiple cloning site with unique restriction enzyme recognition sequences for inserting the gene of interest. (3) A selectable marker (usually antibiotic resistance) so bacteria carrying the plasmid can be identified on selective media.
PCR is a molecular photocopier. Starting from one DNA molecule, PCR can produce billions of copies in a few hours. It is the foundation of nearly every modern DNA technique: forensics, diagnostics, cloning, sequencing, genetic testing.
The Cycle
Each PCR cycle has three temperature steps:
Denaturation (~95°C): heat separates the two DNA strands.
Annealing (~50-65°C): temperature drops, allowing short primers to bind their complementary sequences on each strand.
Extension (~72°C): heat-stable DNA polymerase (Taq) extends each primer 5’ to 3’, doubling the target region.
The three-step PCR cycle. Each cycle doubles the amount of target DNA. After 30 cycles, one molecule becomes more than a billion. Credit: Wikimedia Commons, CC BY-SA
Repeat 25-35 times. One starting molecule becomes 2, then 4, then 8, and so on. After 30 cycles, you have approximately 230 = about 1 billion copies.
Key Ingredients
Template DNA: the starting material (one molecule is enough).
Primers: two short (~20 nt) DNA oligonucleotides that flank the target region. One primer matches the top strand; the other, the bottom strand.
dNTPs: the four nucleotide substrates (dATP, dTTP, dGTP, dCTP).
Taq polymerase: a heat-stable DNA polymerase from the bacterium Thermus aquaticus. It survives 95°C without denaturing, so it does not have to be replaced each cycle.
Buffer + Mg2+: enzymatic cofactors.
Thermal cycler: machine that programs the temperature changes.
Specificity from Primers
PCR’s specificity comes from the primers. Only the region between the two primer binding sites is amplified. If you want to amplify exon 3 of a specific gene, you design primers that flank exon 3. Any other DNA in the sample is ignored because no primers bind there.
Variants
RT-PCR (reverse transcription PCR): starts with RNA. A reverse transcriptase first converts RNA to complementary DNA (cDNA); then standard PCR amplifies the cDNA. Used to measure gene expression.
qPCR (quantitative PCR): real-time PCR with a fluorescent probe that increases signal as amplification proceeds. Allows quantification of starting template. Used heavily in COVID testing.
Multiplex PCR: multiple primer pairs in one reaction to amplify several targets simultaneously.
What are the three temperature steps of a single PCR cycle, and what happens at each?
Click to reveal answer
(1) Denaturation at ~95°C - heat separates the two DNA strands. (2) Annealing at ~55°C - primers hybridize to their complementary sequences. (3) Extension at ~72°C - Taq polymerase extends each primer in the 5' to 3' direction, doubling the target region. Repeat 25-35 times for exponential amplification.
Why is Taq polymerase essential for PCR?
Click to reveal answer
PCR repeatedly heats the reaction to 95°C to denature DNA. Most polymerases would irreversibly denature themselves at this temperature. Taq polymerase, isolated from the thermophilic bacterium Thermus aquaticus, is heat-stable and can survive repeated cycles at 95°C. This lets one dose of enzyme last through all 30+ cycles.
What determines which region of DNA is amplified by PCR?
Click to reveal answer
The two primers. Each primer is a short DNA oligonucleotide complementary to a specific sequence in the template. The region amplified is the stretch between where the two primers bind. Anything outside the primer binding sites is not efficiently amplified. Specificity of PCR comes entirely from primer design.
Gel electrophoresis is the standard way to separate DNA fragments by size. The DNA sample is loaded into wells at one end of a porous gel (usually agarose). An electric field drives the DNA through the gel. Small fragments travel quickly; large fragments travel slowly.
An agarose gel stained with ethidium bromide. DNA bands appear white/orange under UV light. The leftmost lane is a DNA ladder of known fragment sizes for reference. Credit: Wikimedia Commons, CC BY-SA
How It Works
DNA has a uniformly negative charge from its phosphate backbone. In an electric field, it moves toward the positive electrode (anode). The gel matrix has tiny pores that slow fragments based on size:
Small fragments: slip through easily → travel far.
Large fragments: get caught in the mesh → travel less distance.
After the run, fragments are visualized by adding a DNA-binding dye:
Ethidium bromide: intercalates between base pairs and fluoresces under UV light.
SYBR Safe and other modern dyes: less toxic alternatives.
The pattern of bands along each lane is the “gel image” - a visual sorting of the fragments by size.
Reading a Gel
A typical gel image has:
A ladder lane: DNA fragments of known sizes (e.g., 100, 200, 300, 500, 1000, 2000 bp). This is the ruler.
Sample lanes: your DNA samples. Compare bands to the ladder to estimate each fragment’s size.
The wells are at the top (or negative end). Bands near the top are large; bands near the bottom are small.
Agarose vs. Polyacrylamide
Two gel materials are common:
Agarose: best for DNA fragments 100 bp to 50+ kb. Large pores. Lower resolution.
Polyacrylamide: best for small DNA (<500 bp) and for proteins. Smaller pores, higher resolution. Proteins on polyacrylamide (SDS-PAGE) were covered in Chapter 3.
The choice depends on the size range you care about. For routine cloning checks (1-5 kb plasmid fragments), agarose is standard.
Why does DNA migrate toward the positive electrode in gel electrophoresis?
Click to reveal answer
The phosphate backbone of DNA carries many negative charges - one per nucleotide. Under an electric field, the negatively charged DNA is attracted to the positive electrode (anode). Proteins, by contrast, have variable net charges depending on their amino acids and pH, which is why SDS is needed to give proteins uniform negative charge.
In agarose gel electrophoresis, where on the gel do the smallest DNA fragments end up?
Click to reveal answer
Farthest from the wells, closest to the bottom (positive electrode). Small DNA fragments travel more easily through the gel pores, so they migrate faster and end up at the leading edge. Large fragments are stuck in the matrix near the wells. This is the opposite of size exclusion chromatography (big elutes first) because gel electrophoresis physically restricts larger molecules in pores.
How is a band's size estimated from a gel?
Click to reveal answer
By comparison to a ladder - a set of DNA fragments with known sizes run in a parallel lane. The sample band's migration distance is compared to the ladder's bands to estimate the sample's size. More precisely, migration distance is approximately proportional to -log(fragment size), so linear interpolation between adjacent ladder bands gives a good estimate.
A gel tells you how big your fragments are but not which one is the specific molecule you care about. Blotting adds specificity: transfer the gel to a membrane, then probe the membrane with a labeled molecule that binds only your target.
Three main blots, each for a different target molecule type.
The SNoW DRoP Mnemonic
Southern Blot (DNA)
Workflow:
Cut genomic DNA with a restriction enzyme.
Run on a gel to separate fragments by size.
Transfer the DNA onto a membrane (nitrocellulose or nylon).
Incubate the membrane with a labeled DNA probe (radioactive or fluorescent) complementary to the target sequence.
Wash off unbound probe, then detect where the probe stuck.
A single band on the resulting film tells you the target sequence is present in the genome and at what size the restriction enzyme cut it. Southern blotting was historically used to detect gene rearrangements, copy number variations, and specific mutations. It has largely been replaced by PCR and sequencing.
Northern Blot (RNA)
Essentially the same procedure, but starting with RNA. Used to detect and size a specific mRNA. Tells you whether a gene is expressed in a tissue, whether alternative splicing produces multiple transcript sizes, and whether the transcript is degraded. Like Southern, it has been largely supplanted by quantitative methods (qPCR, RNA-seq).
Western Blot (Protein)
Workflow:
Run proteins on SDS-PAGE.
Transfer to a membrane.
Probe with a specific antibody against the target protein.
A secondary antibody (attached to an enzyme or fluorophore) detects the primary antibody.
A chemiluminescent substrate reveals where the primary antibody bound.
Western blotting remains the gold standard for confirming that a specific protein is present (and at what size) in a sample. It is used throughout molecular and clinical biology.
Hybridization Probes vs. Antibodies
The key difference between Southern/Northern and Western blots:
Southern and Northern use complementary nucleic acid probes (DNA or RNA) that base-pair with the target. Specificity comes from sequence complementarity.
Western uses antibodies that recognize a specific protein epitope. Specificity comes from the protein-protein interaction.
What does each of Southern, Northern, and Western blot detect?
Click to reveal answer
Southern detects DNA. Northern detects RNA. Western detects Protein. The mnemonic SNoW DRoP captures this: Southern-DNA, Northern-RNA, Western-Protein.
What is the fundamental probe difference between a Southern blot and a Western blot?
Click to reveal answer
Southern blots use a complementary nucleic acid (DNA or RNA) probe that base-pairs with target DNA. Western blots use antibodies that recognize specific protein epitopes. The recognition principle differs: hybridization vs. protein-protein binding.
Why is Western blot often used as the confirmatory test after a positive ELISA screen for HIV?
Click to reveal answer
ELISA is highly sensitive but can yield false positives. Western blot separates viral proteins by size, then uses the patient's antibodies to confirm that bands at specific sizes (gp120, gp41, p24) are present. Multiple specific bands give high specificity. Screen with the sensitive test (ELISA), confirm with the specific test (Western blot).
DNA sequencing tells you the exact order of nucleotides. The MCAT focuses on Sanger sequencing (the classical chain-termination method) and recognizes that next-generation sequencing (NGS) is the modern high-throughput approach.
Sanger (Chain-Termination) Sequencing
Developed by Frederick Sanger in 1977. The reaction is like a PCR, with one important twist: the reaction mix contains a small amount of dideoxynucleotides (ddNTPs) in addition to the normal dNTPs.
A ddNTP lacks a 3’-OH. When DNA polymerase incorporates a ddNTP, no further nucleotides can be added - the chain terminates at that base.
Sanger sequencing. Each reaction produces fragments of every possible length, each ending in a fluorescently labeled ddNTP. Separation by size reveals which base is at each position. Credit: Wikimedia Commons, CC BY-SA
Workflow
Mix template DNA, primer, DNA polymerase, dNTPs, and a small amount of each ddNTP (each labeled with a different fluorescent color: ddA = green, ddT = red, etc.).
The polymerase extends the primer using dNTPs. Occasionally it incorporates a ddNTP by accident and stops.
The result is a mixture of fragments of every possible length, each ending in a known fluorescent ddNTP.
Run the mixture through capillary electrophoresis (size-separation by size).
A detector reads the fluorescent color of each fragment as it passes. The order of colors = the sequence of bases.
The Chromatogram
The output of a Sanger sequencer is a chromatogram: four colored traces (one per base) with peaks at each position of the sequence. Reading the tallest peak at each position left to right gives the DNA sequence.
Next-Generation Sequencing
Modern NGS platforms (Illumina, PacBio, Nanopore) sequence billions of DNA fragments in parallel. Each platform uses different chemistry, but the general idea is the same: fragment the DNA, attach adapters, amplify, then read out each fragment in parallel. This is how the Human Genome Project went from 13 years and $3 billion (Sanger-based) to 1-2 days and $1000 (modern NGS) for a whole human genome.
You do not need to know specific NGS chemistries for the MCAT. But recognize that modern sequencing is massively parallel, not Sanger’s sequential approach.
What makes a dideoxynucleotide (ddNTP) stop DNA polymerase?
Click to reveal answer
A ddNTP lacks a 3'-OH on its sugar (in addition to lacking the 2'-OH like normal dNTPs). DNA polymerase extends by attacking the 3'-OH of the growing chain with the next nucleotide's 5'-phosphate. No 3'-OH = no attachment point for the next nucleotide = chain terminates. This is the basis of Sanger sequencing.
How is the actual DNA sequence read from a Sanger sequencing reaction?
Click to reveal answer
The reaction produces fragments of every possible length, each ending in a fluorescently labeled ddNTP (one color per base). Capillary electrophoresis separates fragments by size (shortest to longest). A detector reads the color of each fragment in order. The sequence of colors, read from shortest to longest fragment, corresponds to the DNA sequence from the primer outward.
What is the main difference between Sanger sequencing and next-generation sequencing?
Click to reveal answer
Sanger sequences one DNA fragment at a time using chain-termination chemistry. NGS sequences millions of fragments in parallel. NGS is much cheaper and faster per base but produces shorter reads. Modern whole-genome sequencing is almost entirely NGS; Sanger is still used for short, high-accuracy confirmation of specific loci (e.g., confirming a single mutation in a clinical sample).
CRISPR-Cas9 is a programmable DNA-cutting system that researchers have adapted from bacterial immunity. It is the current gold standard for gene editing because it is simple, cheap, and very precise. The 2020 Nobel Prize in Chemistry recognized the development of CRISPR as a gene-editing tool.
The Bacterial Origin
In bacteria, CRISPR (clustered regularly interspaced short palindromic repeats) is an immune system against viruses. Bacteria store short snippets of past viral DNA in a special genomic region. When the virus attacks again, the bacterium transcribes these snippets into guide RNAs that direct a Cas nuclease to cut matching viral DNA.
Molecular biologists hijacked this system. By designing a guide RNA that targets any sequence of interest, Cas9 can be directed to cut any gene in any organism.
CRISPR-Cas9 in action. The guide RNA (blue) base-pairs with the target DNA, and Cas9 (large protein) makes a double-strand cut 3 nucleotides upstream of the PAM (NGG) sequence. Credit: Wikimedia Commons, CC BY-SA
The Two Components
Cas9: an RNA-guided endonuclease that cuts both DNA strands.
Guide RNA (gRNA): a ~100-nucleotide RNA with a 20-nucleotide spacer that base-pairs with the target DNA sequence. The rest of the gRNA binds Cas9.
Target DNA must have a PAM sequence (typically NGG for the common SpCas9) immediately adjacent. The PAM is read by Cas9 and is not part of the gRNA pairing; it prevents Cas9 from cutting bacterial CRISPR arrays (which lack PAMs).
Editing Outcomes
Cas9 makes a blunt double-strand break (DSB) ~3 bp upstream of the PAM. The cell must repair this break. The repair pathway determines the editing outcome:
Non-homologous end joining (NHEJ): the default pathway. The two broken ends are glued back together, often with small insertions or deletions (indels). A frame-shift indel in a coding region typically knocks out the gene. This is how CRISPR produces gene knockouts.
Homology-directed repair (HDR): if a donor DNA template with homologous ends is provided, the cell can use it to repair the break precisely. This enables knock-in of specific edits, corrections of disease mutations, or insertion of tagged sequences.
Applications
Research: creating knockout cell lines and animals. CRISPR has replaced older gene targeting methods almost entirely.
Therapeutics: the first FDA-approved CRISPR therapy, Casgevy (2023), treats sickle cell disease by editing hematopoietic stem cells. Several cancer and genetic disease trials are ongoing.
Diagnostics: CRISPR-based detection systems (SHERLOCK, DETECTR) use Cas13 or Cas12 to detect specific nucleic acid sequences from pathogens.
What are the two core components of CRISPR-Cas9 gene editing?
Click to reveal answer
Cas9 (an RNA-guided endonuclease) and a guide RNA. The guide RNA has a 20-nt spacer that base-pairs with the target DNA and a scaffold region that binds Cas9. The gRNA directs Cas9 to the target site, where Cas9 makes a double-strand break, provided an adjacent PAM sequence (usually NGG) is present.
How does CRISPR produce a gene knockout vs. a precise gene correction?
Click to reveal answer
The difference is the repair pathway. Without a donor template, the cell repairs the DSB by NHEJ, introducing small indels that often disrupt the reading frame - knockout. With a donor DNA template flanked by homology to the cut site, the cell can use HDR to install a precise edit from the template - knock-in or gene correction.
What is a PAM sequence and why does it matter for CRISPR?
Click to reveal answer
PAM (protospacer adjacent motif) is a short DNA sequence adjacent to the target site, most commonly NGG for the widely used SpCas9. Cas9 binds the PAM before checking the gRNA-target match. If there is no valid PAM, Cas9 does not cut. The PAM requirement protects bacterial CRISPR arrays (which lack PAMs) and constrains which sites in a target genome can be edited.