Difficulty: Intermediate | Prerequisites: Lessons 1–3 notes (Central Dogma, nucleotide structure, protein basics)
Lessons 4 through 6 move from the architecture of chromosomes and chromatin (Lesson 4), through the mechanics of DNA replication and repair (Lesson 5), to how genes are transcribed into RNA and translated into protein (Lesson 6). Together they explain how cells faithfully copy their genomes, fix mistakes, and convert genetic information into functional molecules. You need the nucleotide chemistry from Lesson 1 and the enzyme concepts from Lesson 3 to follow the machinery here.
DNA is packaged into chromatin by wrapping around histone octamers (nucleosomes), and the pattern of histone modifications controls which genes are accessible. DNA replication proceeds bidirectionally from origins, with distinct mechanisms for leading and lagging strands, and multiple repair pathways correct errors and damage. Transcription produces pre-mRNA that is capped, spliced, and polyadenylated before export, and translation occurs on ribosomes through a precise initiation–elongation–termination cycle.
Deoxyribonucleotide
A DNA monomer consisting of a base, a deoxyribose sugar, and a phosphate group.
Deoxyribonucleoside
A base attached to a deoxyribose sugar, without the phosphate group. Think of it as a nucleotide minus the phosphate.
Chromatin
Linear DNA duplex molecules bound to proteins (histones and non-histone proteins). This is the form DNA takes inside the nucleus.
Euchromatin
Chromatin that remains in an open, decondensed state during interphase. Contains most of the actively expressed genes. Associated with histone acetylation.
Heterochromatin
Chromatin that stays condensed during interphase. Genes within it are largely silenced. Associated with histone methylation.
Nucleosome
The fundamental unit of chromatin packaging: a histone octamer (2 copies each of H2A, H2B, H3, H4) wrapped by 147 base pairs of DNA. Think of it as a bead on a string, where the bead is the histone core and the string is the DNA.
Histone code
The combinatorial pattern of covalent modifications on histone tails (methylation, acetylation, phosphorylation) that signals whether a region of chromatin should be active or silent. Specific proteins read, write, and erase these marks.
Epigenetics
The regulation of chromatin structure and gene activity through heritable modifications that do not change the DNA sequence itself. Histone modifications and DNA methylation are the primary mechanisms.
Intron
A non-coding sequence within a gene that is transcribed into pre-mRNA but removed by splicing before translation.
Exon
A coding sequence within a gene that is retained in the mature mRNA and encodes part of the protein's amino acid sequence.
DNA polymerase
The enzyme that synthesises new DNA strands in the 5ʹ → 3ʹ direction by adding nucleotides complementary to a template strand. Also has 3ʹ → 5ʹ exonuclease activity for proofreading.
DNA primase
An enzyme that synthesises short RNA primers required to initiate DNA replication. Only one primer is needed for the leading strand; multiple primers are needed for the lagging strand.
DNA ligase
An enzyme that joins DNA fragments by forming a phosphodiester bond between adjacent nucleotides. Consumes ATP and releases AMP.
Topoisomerase
An enzyme that prevents DNA tangling during replication by temporarily breaking the phosphodiester backbone. Topo I makes single-strand breaks (no ATP needed); Topo II makes double-strand breaks (ATP required).
Telomerase
An enzyme that extends the 3ʹ end of telomeres using an internal RNA template, solving the end-replication problem. Uses reverse transcription (RNA → DNA).
Base excision repair (BER)
A repair pathway where DNA glycosylase recognises and removes an altered base, and AP endonuclease then removes the remaining sugar and phosphate. The gap is filled by polymerase and sealed by ligase.
Nucleotide excision repair (NER)
A repair pathway for bulky lesions (e.g. pyrimidine dimers). A protein complex detects the distortion, endonucleases cut the backbone on both sides, helicase removes the damaged segment, and polymerase and ligase fill the gap.
Nonhomologous end joining (NHEJ)
A double-strand break repair pathway that directly joins broken DNA ends after processing. It is fast but mutagenic because it often causes small deletions.
Homologous recombination (HR)
A double-strand break repair pathway that uses the sister chromatid (made during S phase) as a template to accurately restore the original sequence.
RNA polymerase II (Pol II)
The eukaryotic RNA polymerase responsible for transcribing all protein-coding genes into mRNA, as well as snoRNA and most snRNA genes.
Spliceosome
A large ribonucleoprotein complex (containing snRNPs and accessory proteins) that removes introns from pre-mRNA and joins exons together.
Alternative splicing
The process by which the same pre-mRNA can be spliced in different ways in different cell types, producing tissue-specific protein isoforms from a single gene.
Ribosome (80S, eukaryotic)
The molecular machine that translates mRNA into protein. Composed of a large subunit (60S) and a small subunit (40S), with four RNA binding sites: one for mRNA and three for tRNA (A site, P site, E site).
Polysome (polyribosome)
An mRNA molecule being translated simultaneously by multiple ribosomes, increasing the rate of protein production from a single transcript.
DNA Ends and Chromosome Organisation
5ʹ end of DNA: phosphate group; 3ʹ end: hydroxyl group of sugar
Human chromosomes: 44 autosomes (22 pairs) + 2 sex chromosomes (XX or XY)
Banding patterns on chromosomes reflect local organisation of condensed chromatin
Chromosome Replication and Cell Division
Before S phase: 1 interphase chromosome = 1 DNA molecule (1 duplex, 1 centromere, 2 telomeres)
After S phase: 1 interphase chromosome = 2 DNA molecules (2 duplexes, 2 centromeres, 4 telomeres)
During mitosis: 1 mitotic chromosome = 2 sister chromatids held together by centromeres
Chromatin and Nucleosomes
Euchromatin: open conformation, acetylated histones, most genes expressed
Heterochromatin: condensed, methylated histones, most genes silenced
Nucleolus: site where rRNA genes are transcribed
Nucleosome = histone octamer (2 × H2A, H2B, H3, H4) + 147 bp of DNA
Linker DNA between nucleosomes binds histone H1
Histone Modifications and the Histone Code
Histone N-terminal tails extend from the nucleosome and interact with DNA, other proteins, and are targets for covalent modifications
Lysine methylation: preserves positive charge (1–3 methyl groups possible)
Lysine acetylation: neutralises positive charge, carried out by HATs (histone acetyl transferases) and reversed by HDACs (histone deacetylase complexes)
Serine phosphorylation: adds negative charge (by kinases)
The pattern of modifications constitutes a "histone code" that is recognised, copied, and erased by specific proteins
Epigenetics: heritable regulation of chromatin structure that influences gene expression without altering the DNA sequence
Genome Structure
Expressed genes sit on exposed chromatin loops during interphase; chromosomal loops can change position in the nucleus as gene expression changes
Each chromosome occupies a distinct 3D region in the nucleus (nuclear territory)
Repeated sequences: LINEs, SINEs, transposons (mobile genetic elements)
Only ~1.5% of the human genome consists of protein-coding exons
Introns: intervening non-coding sequences; Exons: sequences encoding amino acids
KAT6B Mutations (Clinical Application)
Mutations in the KAT6B gene produce a truncated histone acetyltransferase (premature stop codon) that functions abnormally during early development
Associated conditions: Say-Barber-Biesecker-Young-Simpson Syndrome (SBBYSS) and Genitopatellar Syndrome (GTPTS)
Replication Machinery
DNA replication proceeds 5ʹ → 3ʹ on the leading strand
DNA polymerase: adds nucleotides to the template strand; has two catalytic sites
P site for polymerase activity (5ʹ → 3ʹ synthesis)
E site for exonuclease proofreading (3ʹ → 5ʹ editing)
DNA primase: synthesises short RNA primers; one primer for the leading strand, many for the lagging strand
DNA polymerase also removes RNA primers after use
DNA ligase: seals gaps by forming phosphodiester bonds (consumes ATP, releases AMP)
Topoisomerases
Topo I: single-strand break, allows rotation to relieve tension, reseals without ATP
Topo II: double-strand break, allows one duplex to pass through another, requires ATP
Replication Origins and Bubbles
Replication origins are AT-rich (A–T has only 2 H-bonds, easier to separate than G–C with 3)
Origin recognition complex (ORC) binds origins in G1
Replication is bidirectional: two forks move away from each origin
Leading strand synthesis is continuous; lagging strand synthesis is discontinuous (Okazaki fragments)
Nucleosomes are reassembled on new DNA behind the replication fork, with the help of histone chaperones NAP1 and CAF1
Origin Licensing
Ensures each origin fires only once per S phase:
G1: ORC binds origins
Late G1: pre-replicative complex forms
Early S: DNA synthesis begins, ORC is phosphorylated and inactivated
G2: ORC remains phosphorylated
Next G1: ORC is dephosphorylated, ready for the next round
Telomerase
Solves the end-replication problem by extending the 3ʹ end of the template strand at telomeres
Uses its own RNA component as a template (reverse transcription)
DNA Damage and Mutation
Deamination: removal of a functional group from a base (can cause miscoding)
Depurination: removal of an entire purine base (creates an abasic site)
Pyrimidine dimers (most commonly thymine dimers): formed by UV light; block DNA polymerase, are mutagenic, and can lead to base pair additions or deletions
Repair Pathways
BER: DNA glycosylase removes the damaged base → AP endonuclease removes the sugar-phosphate → polymerase and ligase fill and seal the gap
NER: detects bulky distortions (e.g. pyrimidine dimers) → endonuclease cuts on both sides → helicase removes the damaged segment → polymerase and ligase repair the gap
NER is coupled to transcription by RNA polymerase, prioritising repair of actively expressed genes
Double-Strand Break Repair
NHEJ: fast but error-prone; DNA ends are processed and joined directly, often losing a few nucleotides (mutagenic)
Homologous recombination (HR): accurate; uses the sister chromatid from S phase as a template
DSB is processed to single-stranded ends → one strand invades the sister chromatid → extended using the sister as template → released and re-annealed
Can also repair broken replication forks
HR in Meiosis
Crossing over and gene conversion produce unique chromosomes
HR occurs between homologous chromosomes (not sister chromatids as in mitotic repair)
Transposons
Mobile genetic elements that encode transposase; do not require DNA homology for movement
BRCA1 and BRCA2 (Clinical Application)
Tumour suppressor genes; their protein products are required for homologous recombination
Mutations in BRCA1/BRCA2 impair HR, leading to persistent DNA damage accumulation
Most common cause of hereditary breast and ovarian cancer
RNA Basics
Ribose in RNA has a hydroxyl group at the 2ʹ position (DNA has hydrogen)
Types of RNA: mRNA, rRNA, tRNA, snRNA (splicing of pre-mRNA), snoRNA (rRNA processing and modification)
RNAs form non-conventional base pairs to stabilise 3D structure
RNA Polymerases
RNA polymerase starts transcription without a primer and has proofreading capability
Pol I: transcribes 5.8S, 18S, 28S rRNA genes
Pol II: transcribes all protein-coding genes, snoRNA genes, most snRNA genes
Pol III: transcribes tRNA genes, 5S rRNA genes, some snRNA genes
Larger S value = larger rRNA molecule
Pol II Transcription
Eukaryotic promoters recognised by Pol II contain a TATA consensus sequence
TATA box is recognised by TFIID (transcription factor IID)
General transcription factor (TF) complexes regulate initiation, elongation, and termination
Enhancers: DNA sequences that increase promoter activity (can be distant from the gene)
mRNA Processing
Pre-mRNA is processed co-transcriptionally:
5ʹ cap is added
Introns are removed by splicing
3ʹ poly-A tail is added
The spliceosome (snRNPs + accessory proteins) carries out splicing
snRNAs recognise splice site consensus sequences and catalyse splicing reactions
Alternative splicing: the same pre-mRNA transcript produces different protein isoforms in different cell types by including or excluding specific exons
RNA Export
In the nucleus, mRNAs associate with cap binding complex (CBC), SR proteins, hnRNP proteins, and poly-A tail binding proteins
Export receptors carry mRNAs through nuclear pore complexes to the cytoplasm
In the cytoplasm, mRNAs bind eukaryotic initiation factors (eIFs) for translation
rRNA Processing
5.8S, 18S, and 28S rRNAs are synthesised together as a 45S precursor by Pol I in the nucleolus (5S is transcribed separately by Pol III)
Processing and chemical modifications guided by snoRNAs in snoRNPs
tRNA and Aminoacyl-tRNA Synthetases
Each tRNA is specific for one amino acid, which is "charged" onto its 3ʹ end
Aminoacyl-tRNA synthetases use ATP to link specific amino acids to their corresponding tRNAs
Translation: The Ribosome
Eukaryotic 80S ribosome = 60S (large) + 40S (small)
4 RNA binding sites: 1 for mRNA + 3 for tRNA
A site (aminoacyl): incoming charged tRNA binds here
P site (peptidyl): holds tRNA attached to the growing polypeptide
E site (exit): deacylated tRNA exits here
Translation: Initiation
Small ribosomal subunit binds mRNA and scans for the start codon (AUG)
Initiator tRNA (carrying methionine) binds the P site
Large subunit joins to form the complete ribosome
Translation: Elongation
Charged tRNA with matching anticodon enters the A site
Peptidyl transferase (in the large subunit) catalyses peptide bond formation, transferring the growing chain from P-site tRNA to A-site amino acid
Energy comes from the high-energy bond between the amino acid and tRNA in the P site
Ribosome translocates one codon: large subunit moves first, then small subunit
Empty tRNA is ejected from the E site
Translation: Termination
Stop codons (UAA, UAG, UGA) are not recognised by any tRNA
Release factors bind the ribosome when a stop codon enters the A site
Peptidyl transferase adds a water molecule to the peptidyl-tRNA in the P site, releasing the completed protein with a carboxylic acid group at its C-terminus
Translation accuracy: ~99.9%, maintained by two proofreading steps
Protein Folding After Translation
Folding begins as the nascent polypeptide exits the ribosome (N-terminus folds first)
Post-translational modifications and assembly with cofactors or other proteins may follow
Foldase enzymes help fold polypeptides (can hydrolyse ATP)
Protein chaperones (e.g. Hsp70) promote correct 2° and 3° structures by hydrolysing ATP
Regulation of Gene Expression
Six control points: transcription, RNA processing, export and localisation, translation, RNA degradation, protein modification and degradation
Antibiotics and Translation (Clinical Application)
Many antibiotics exploit structural differences between bacterial and eukaryotic ribosomes
Tetracycline: blocks aminoacyl-tRNA binding to the A site of bacterial ribosomes
Chloramphenicol: blocks the peptidyl transferase reaction on bacterial ribosomes
Nucleosome: histone octamer + 147 bp of DNA
Human genome: ~3.2 billion bp, ~21,000 protein-coding genes, ~1.5% protein-coding exons, ~50% repetitive elements
Glycolysis net yield: 2 ATP (from Lesson 2, but frequently tested alongside metabolism of the cell)
Krebs cycle per turn: 3 NADH, 1 GTP, 1 FADH₂, 2 CO₂
Translation accuracy: 99.9%
BRCA1/BRCA2 genetic testing is now routine in clinical oncology. Patients with mutations in these HR-repair genes face elevated risks of breast and ovarian cancer, and targeted therapies (PARP inhibitors) exploit the cells' inability to repair DNA.
Understanding NER explains why patients with xeroderma pigmentosum (defective NER) are extremely sensitive to UV light and develop skin cancers at a very young age.
Antibiotics like tetracycline and chloramphenicol work because bacterial ribosomes are structurally different from human ribosomes, allowing selective toxicity.
Students often think histones merely "package" DNA for compactness. They also regulate gene expression: the histone code determines which genes are active or silenced.
Students sometimes mix up euchromatin and heterochromatin. Remember: euchromatin = "eu" (true/good) = expressed, open, acetylated. Heterochromatin = condensed, silenced, methylated.
Students frequently confuse leading and lagging strand synthesis. The leading strand is synthesised continuously (one primer); the lagging strand is synthesised in discontinuous Okazaki fragments (many primers).
Students often think NHEJ and HR are interchangeable. NHEJ is fast but mutagenic (causes deletions); HR is accurate but requires a sister chromatid template (available only after S phase).
⚠️ Know the composition of a nucleosome (octamer of H2A, H2B, H3, H4 + 147 bp DNA + linker DNA with H1).
⚠️ Be able to compare euchromatin vs heterochromatin in terms of conformation, histone modifications, and gene activity.
⚠️ Understand origin licensing and how the cell ensures each origin fires only once per S phase.
⚠️ Know the difference between BER (single damaged base) and NER (bulky lesion/pyrimidine dimer).
⚠️ Be able to compare NHEJ (fast, mutagenic) with HR (accurate, needs sister chromatid).
⚠️ Understand the three eukaryotic RNA polymerases and what each transcribes.
⚠️ Be able to walk through translation step by step: initiation, elongation (A site, P site, E site), termination.
⚠️ Know how BRCA1/BRCA2 mutations lead to cancer through impaired homologous recombination.
True or False: Heterochromatin is associated with histone acetylation and active gene expression.
Fill in the blank: Replication origins are rich in ______ base pairs because they require only 2 hydrogen bonds to separate.
True or False: Topoisomerase I requires ATP to reseal the backbone after making a single-strand break.
Fill in the blank: In translation, peptide bond formation is catalysed by ______ activity in the large ribosomal subunit.
True or False: Alternative splicing allows one gene to produce multiple different protein isoforms.
Q: What is a nucleosome, and what is its composition?
A: A nucleosome is the fundamental unit of chromatin. It consists of a histone octamer (two copies each of H2A, H2B, H3, and H4) around which 147 base pairs of DNA are wrapped. Linker DNA between nucleosomes is associated with histone H1.
Q: How does histone acetylation affect gene expression?
A: Acetylation of lysine residues on histone tails neutralises their positive charge, weakening the interaction between histones and negatively charged DNA. This opens the chromatin (euchromatin), making genes more accessible for transcription.
Q: Why are replication origins rich in AT base pairs?
A: A–T base pairs are held together by only two hydrogen bonds (compared to three for G–C), making them easier to separate. This facilitates the initial strand separation needed to begin replication.
Q: Compare the BER and NER repair pathways.
A: BER handles single damaged or altered bases: DNA glycosylase removes the base, AP endonuclease removes the sugar-phosphate, then polymerase and ligase fill the gap. NER handles bulky lesions such as pyrimidine dimers: a protein complex detects the distortion, endonucleases cut on both sides, helicase removes the damaged strand segment, and polymerase and ligase repair the gap.
Q: Why is NHEJ considered mutagenic while HR is considered accurate?
A: NHEJ processes the broken DNA ends before joining them, which typically results in small deletions of nucleotide sequence. HR uses the intact sister chromatid as a template to restore the original sequence precisely.
Q: Walk through the elongation step of translation.
A: A charged tRNA with the correct anticodon enters the A site. Peptidyl transferase in the large subunit transfers the growing polypeptide from the P-site tRNA to the amino acid on the A-site tRNA, forming a new peptide bond. The ribosome then translocates one codon: the large subunit shifts first, then the small subunit. The now-empty tRNA exits from the E site, and the next codon is exposed in the A site.
Q: How do BRCA1/BRCA2 mutations contribute to cancer?
A: BRCA1 and BRCA2 proteins are essential for homologous recombination repair of double-strand breaks. When both copies are mutated and non-functional, the cell cannot accurately repair DSBs. This leads to accumulating DNA damage over time, increasing the risk of oncogenic mutations and hereditary breast and ovarian cancer.
Chromatin structure and histone modifications (Lesson 4) connect to the regulation of gene expression covered at the end of Lesson 6. The histone code determines which promoters Pol II can access.
DNA repair (Lesson 5) ties into the clinical discussion of BRCA mutations and the broader theme of how cells maintain genomic integrity across the cell cycle.
The translation machinery (Lesson 6) relies on the protein-folding principles from Lesson 3: once a polypeptide exits the ribosome, its amino acid sequence (primary structure) dictates how it folds into secondary, tertiary, and quaternary structures.
deoxyribonucleotide, deoxyribonucleoside, chromatin, euchromatin, heterochromatin, nucleosome, histone octamer, H2A, H2B, H3, H4, histone H1, linker DNA, histone code, histone acetylation, HAT, HDAC, histone methylation, epigenetics, chromosome, centromere, telomere, sister chromatid, intron, exon, LINE, SINE, transposon, DNA polymerase, DNA primase, DNA ligase, topoisomerase, Okazaki fragment, leading strand, lagging strand, replication origin, ORC, origin licensing, telomerase, reverse transcription, deamination, depurination, pyrimidine dimer, thymine dimer, UV damage, base excision repair, BER, nucleotide excision repair, NER, AP endonuclease, DNA glycosylase, double-strand break, NHEJ, homologous recombination, HR, BRCA1, BRCA2, tumour suppressor, RNA polymerase, Pol I, Pol II, Pol III, TATA box, TFIID, transcription factor, enhancer, pre-mRNA, 5 prime cap, poly-A tail, spliceosome, snRNP, snRNA, alternative splicing, mRNA export, nuclear pore complex, rRNA, tRNA, aminoacyl-tRNA synthetase, ribosome, 80S, 60S, 40S, A site, P site, E site, peptidyl transferase, translation initiation, elongation, termination, stop codon, release factor, polysome, protein folding, chaperone, Hsp70, foldase, tetracycline, chloramphenicol, KAT6B, SBBYSS