Source: Cell Biology, University of Florida | Chapter 7
Tags: central dogma, transcription, translation, mRNA, tRNA, rRNA, RNA polymerase, ribosome, codon, anticodon, gene expression, protein synthesis, BIOL 101
Difficulty: Intermediate | Prerequisites: Basic DNA structure, nucleotide base pairing (Chapter 5/6 material)
This chapter is the backbone of molecular biology. It explains how the information stored in DNA is used to build proteins, the molecules that do most of the work inside a cell. The process has two main stages: transcription (copying DNA into RNA) and translation (reading RNA to assemble a protein). If you understand this chapter well, nearly every topic that follows in the course, from gene regulation to cell signalling, will make considerably more sense. You should already be comfortable with DNA structure, complementary base pairing, and the difference between nucleotides and amino acids.
DNA is copied into RNA by transcription, and RNA is read by ribosomes to build proteins during translation. The genetic code uses three-nucleotide codons to specify amino acids, and transfer RNAs act as adaptors that match each codon to the right amino acid. In eukaryotes, the RNA transcript must be processed (capped, spliced, and polyadenylated) before it leaves the nucleus for translation.
Central dogma
The principle that genetic information flows from DNA to RNA to protein. DNA serves as the template for RNA synthesis, and RNA directs protein synthesis. In simple terms, this means: DNA is the master blueprint, RNA is the working copy, and protein is the finished product.
Transcription
The process of copying a gene's nucleotide sequence from DNA into a complementary RNA molecule, carried out by RNA polymerase. Think of it as: photocopying one page out of a reference book so the original never leaves the library.
Translation
The process by which the nucleotide sequence of an mRNA is decoded by ribosomes and tRNAs to produce a polypeptide chain (protein). Think of it as: reading a sentence written in one language and converting it, word by word, into another language.
mRNA (messenger RNA)
The RNA molecule that carries the protein-coding information from a gene to the ribosome. In simple terms, this means: the portable copy of the gene that the ribosome reads.
RNA polymerase
The enzyme that synthesises RNA by reading the DNA template strand and adding complementary ribonucleotides in the 5' to 3' direction.
Template strand
The DNA strand that RNA polymerase reads (3' to 5') to produce a complementary RNA molecule. Also called the antisense strand.
Coding strand (non-template strand)
The DNA strand that has the same sequence as the RNA product (with T instead of U). It is not read by RNA polymerase during transcription.
Promoter
A specific DNA sequence upstream of a gene that signals RNA polymerase where to bind and begin transcription. Promoter sequences are not themselves transcribed.
Terminator
A DNA sequence that signals RNA polymerase to stop transcription and release the RNA transcript. Terminator sequences are transcribed into the RNA.
Sigma factor
A protein subunit in bacteria that associates with RNA polymerase and helps it recognise and bind to the promoter. It scans and binds at the -35 position on the outside of the DNA helix.
General transcription factors
A set of accessory proteins required by eukaryotic RNA polymerase II to initiate transcription. They assemble at the promoter and help open the DNA helix. They are the eukaryotic equivalent of the bacterial sigma factor.
TATA box
A short segment of DNA in eukaryotic promoters composed primarily of T and A nucleotides, recognised by the general transcription factor TFIID (specifically its TBP subunit). Acts as a landmark for assembling the transcription initiation complex.
TFIID / TBP (TATA-binding protein)
TFIID is the general transcription factor that first binds the TATA box. TBP is the subunit of TFIID that makes direct contact with the DNA, causing a local distortion in the double helix.
TFIIH
A general transcription factor with kinase activity. It pries apart the DNA helix at the transcription start site (requiring ATP) and phosphorylates the tail of RNA polymerase II, signalling it to begin elongation.
RNA polymerase I, II, III
Three distinct RNA polymerases in eukaryotes. Pol I transcribes most rRNA genes. Pol II transcribes all protein-coding genes, miRNA genes, and some noncoding RNA genes (the one you will encounter most). Pol III transcribes tRNA genes, the 5S rRNA gene, and other small RNAs.
Pre-mRNA processing (capping, splicing, polyadenylation)
The three co-transcriptional modifications made to eukaryotic mRNA before it is exported from the nucleus. Capping adds a modified guanine to the 5' end. Splicing removes introns and joins exons. Polyadenylation adds a poly-A tail to the 3' end.
5' cap
An atypical guanine nucleotide with a methyl group, added to the 5' end of the mRNA after approximately 25 nucleotides have been synthesised. Increases RNA stability, facilitates nuclear export, and marks the molecule as mRNA.
Poly-A tail
A stretch of roughly 150 to 250 adenine nucleotides added to the 3' end of the mRNA after it is trimmed by a specific enzyme. Increases stability and aids nuclear export.
Introns
Non-coding sequences that interrupt eukaryotic genes. They are transcribed into pre-mRNA but removed by splicing before translation. Introns range from 1 to 10,000 nucleotides and are typically longer than exons.
Exons
The expressed (coding) sequences of a gene that remain in the mature mRNA after splicing and are translated into protein.
RNA splicing
The process by which introns are removed from pre-mRNA and exons are joined together, carried out largely by small nuclear ribonucleoproteins (snRNPs). The excised intron forms a lariat structure.
snRNPs (small nuclear ribonucleoproteins)
Complexes of RNA and protein that recognise splice-site sequences in pre-mRNA through complementary base pairing, and catalyse the splicing reaction.
Codon
A sequence of three consecutive nucleotides in mRNA that specifies a particular amino acid (or a stop signal). There are 64 possible codons: 61 code for amino acids and 3 are stop codons.
Start codon (AUG)
The codon that signals the beginning of translation. It codes for methionine and sets the correct reading frame.
Stop codons (UAG, UAA, UGA)
Three codons that do not code for any amino acid. They signal the ribosome to terminate translation. Recognised by release factors, not by tRNAs.
Reading frame
The way in which the nucleotide sequence of an mRNA is partitioned into successive, non-overlapping triplets (codons). Only one of three possible reading frames is correct for a given mRNA, determined by where decoding begins.
tRNA (transfer RNA)
Small RNA molecules that serve as adaptors during translation. Each tRNA carries a specific amino acid at its 3' end and has an anticodon that base-pairs with the complementary codon on the mRNA.
Anticodon
A set of three consecutive nucleotides on a tRNA molecule that base-pairs with the complementary codon on the mRNA.
Wobble base pairing
The relaxed base-pairing rules at the third position of a codon, allowing some tRNAs to recognise more than one codon for the same amino acid. This explains why many synonymous codons differ only in their third nucleotide.
Aminoacyl-tRNA synthetase
An enzyme that covalently attaches the correct amino acid to its corresponding tRNA, using ATP. There are 20 different synthetases, one per amino acid. This charging step is essential for accurate translation.
Ribosome
A large molecular machine composed of ribosomal RNAs and proteins, organised into a large and a small subunit. The ribosome binds mRNA, positions tRNAs, and catalyses peptide bond formation during translation.
A site, P site, E site
The three tRNA-binding sites on the ribosome. A (aminoacyl/acceptor) site: where the incoming charged tRNA binds. P (peptidyl) site: where the tRNA holding the growing polypeptide chain sits. E (exit) site: where the spent tRNA is released.
Ribozyme
An RNA molecule with enzymatic (catalytic) activity. The ribosome itself is a ribozyme: the rRNA, not the protein components, catalyses peptide bond formation.
Polyribosome (polysome)
Multiple ribosomes translating a single mRNA molecule simultaneously, enabling rapid production of many copies of the same protein.
Release factors
Proteins that bind to a stop codon in the A site, causing addition of water (rather than an amino acid) to the polypeptide chain. This frees the completed protein and causes the ribosome to dissociate.
Proteasome
A large protein complex in the cytosol and nucleus that degrades proteins marked with ubiquitin. Consists of a central proteolytic cylinder capped by regulatory complexes. Uses ATP to unfold and thread target proteins into its interior.
Ubiquitin
A small protein covalently attached to other proteins to mark them for degradation by the proteasome.
Gene expression
The process by which information encoded in a DNA sequence is converted into a functional product, such as RNA or protein.
Information flows from DNA to RNA (transcription) and from RNA to protein (translation)
DNA does not make protein directly; it acts as a manager, delegating tasks to RNA intermediaries
A gene is a segment of DNA whose nucleotide sequence is copied into RNA
Cells can control the rate of transcription and translation of each gene independently, allowing fine-tuned regulation of protein levels
DNA: double-stranded, uses deoxyribose sugar, contains thymine (T), primarily stores genetic information
RNA: single-stranded, uses ribose sugar, contains uracil (U) in place of thymine, performs structural, regulatory, and catalytic roles
Base pairing in DNA: A pairs with T. In RNA: A pairs with U
RNA molecules can fold into specific three-dimensional structures held together by internal base pairing
RNA polymerase opens a small portion of the DNA double helix, reads the template strand (3' to 5'), and synthesises the RNA chain in the 5' to 3' direction
Ribonucleotides are added one at a time by complementary base pairing with the template strand
The incoming ribonucleoside triphosphate provides the energy for the reaction (its high-energy phosphate bonds are cleaved)
The newly made RNA strand does not remain hydrogen-bonded to the template; it peels away, and the DNA helix reforms behind the polymerase
RNA transcripts are therefore single-stranded and much shorter than the full DNA molecule (typically a few thousand bases)
RNA polymerase uses ribonucleoside triphosphates as substrates (not deoxyribonucleotides)
RNA polymerase can start a new RNA chain without a primer
RNA polymerase does not proofread accurately; error rate is approximately 1 in 10,000 nucleotides
The RNA product is single-stranded and is displaced from the template
Promoter: a sequence upstream of the gene that tells RNA polymerase where to bind. The promoter is not transcribed. Its polarity (asymmetric 5' to 3' sequence) determines which strand serves as the template and the direction of transcription
Terminator: a sequence that tells polymerase to stop. Terminator sequences are transcribed into the RNA
On a single chromosome, different genes can use different strands as the template, depending on promoter orientation
Bacteria have one RNA polymerase
The sigma factor associates with RNA polymerase and recognises the promoter
Sigma factor first scans and binds at the -35 region from the outside of the double helix (no strand separation needed)
It then opens the helix and examines the -10 region on individual strands
After initiation, sigma factor is released and polymerase clamps onto DNA for elongation
Elongation continues until the polymerase hits the terminator sequence; the enzyme then stops and releases both the DNA template and the RNA transcript
Eukaryotes have three RNA polymerases (I, II, III), each transcribing different classes of genes
RNA Pol II is the most important for this course: it transcribes all protein-coding genes
Eukaryotic RNA polymerase II cannot initiate transcription on its own; it requires a large set of general transcription factors
Assembly at the promoter:
TFIID (via its TBP subunit) binds the TATA box, distorting the DNA
This enables sequential binding of TFIIB, other general transcription factors, and RNA Pol II, forming the transcription initiation complex
TFIIH uses ATP to pry open the helix at the start site, then phosphorylates the tail of RNA Pol II (kinase activity), signalling transcription to begin
Once elongation starts, most general transcription factors (except TFIID) dissociate and are recycled
Elongation factors load onto the polymerase to help it move through nucleosome-packed chromatin
When Pol II finishes a gene, it releases from the DNA, phosphatases strip the phosphates from its tail, and only the dephosphorylated form can reinitiate at a new promoter
mRNA (messenger RNA): carries protein-coding information
rRNA (ribosomal RNA): forms the structural and catalytic core of the ribosome
tRNA (transfer RNA): adaptor molecules that match amino acids to mRNA codons during translation
miRNA (microRNA): regulates gene expression
siRNA (small interfering RNA): provides protection from viruses and transposable elements
lncRNA (long noncoding RNA): acts as scaffolds, diverse regulatory functions
Other noncoding RNAs: involved in splicing, gene regulation, telomere maintenance
Eukaryotic mRNA undergoes three processing steps, often co-transcriptionally (enzymes ride on the phosphorylated tail of RNA Pol II):
5' capping: a methylated guanine is added to the 5' end after about 25 nucleotides have been synthesised. Functions: increases stability, facilitates nuclear export, marks it as mRNA
RNA splicing: introns are removed and exons are stitched together
Special nucleotide sequences at or near intron-exon boundaries act as cues
snRNPs recognise these sequences by complementary base pairing, then cut out the intron as a lariat structure and ligate the exons
Bacterial genes are generally uninterrupted (no introns); eukaryotic protein-coding genes typically contain introns that are longer than the exons
3' polyadenylation: the 3' end is trimmed at a specific sequence, then an enzyme adds approximately 150 to 250 adenine nucleotides (poly-A tail). Functions: increases stability, aids export
Processing occurs simultaneously with transcription, so a full-length unspliced transcript rarely exists in the cell
Only correctly processed mRNAs are exported from the nucleus to the cytosol through nuclear pores
RNA polymerases and processing proteins form loose molecular aggregates called biomolecular condensates
These act as "factories" that bring together the machinery and the genes to be expressed
They are highly organised nuclear subdomains
mRNA molecules are eventually degraded in the cytosol by ribonucleases (RNases)
Lifespan is variable: 30 minutes to less than 10 hours, depending on the sequence and the cell type
In bacteria, most mRNAs are degraded rapidly
mRNA lifespan is one mechanism for controlling how much protein is produced
Transcription is relatively straightforward because DNA and RNA use similar chemical "alphabets" (nucleotides)
Translation is fundamentally different: the information must be converted from the language of nucleotides (4 types) to the language of amino acids (20 types)
The genetic code is the set of rules for this conversion
mRNA is read in groups of three nucleotides called codons
4 x 4 x 4 = 64 possible codons
61 codons specify amino acids; 3 are stop signals (UAG, UAA, UGA)
The code is redundant (degenerate): most amino acids are specified by more than one codon
AUG is the universal start codon and codes for methionine
Each mRNA has three possible reading frames, but only one is correct; the reading frame is set by where decoding begins (the first AUG)
tRNA molecules bridge the gap between mRNA codons and amino acids
Structure:
Folds into a cloverleaf shape with four double-helical segments
Two critical unpaired regions: the anticodon (base-pairs with the mRNA codon) and the 3' end (where the amino acid is covalently attached)
Wobble base pairing at the third codon position allows some tRNAs to recognise multiple codons for the same amino acid
Aminoacyl-tRNA synthetases (20 in total, one per amino acid) charge each tRNA with the correct amino acid, using ATP
The enzyme recognises both the anticodon loop and the amino acid-accepting arm of the tRNA
The reaction creates a high-energy bond between the tRNA and the amino acid; the energy stored in this bond is later used to form the peptide bond on the ribosome
A large complex of rRNAs and proteins, comprising a small subunit and a large subunit
Eukaryotic ribosome: approximately 82 proteins + 4 rRNA molecules (MW approximately 4,200,000)
The small subunit matches tRNAs to mRNA codons; the large subunit catalyses peptide bond formation
The ribosome has three tRNA-binding sites: A (aminoacyl), P (peptidyl), E (exit)
The ribosome is a ribozyme: rRNA (not protein) is responsible for its catalytic activity and overall structure; proteins mainly stabilise the RNA core
The A, P, E sites and the catalytic site for peptide bond formation are all formed by rRNAs
Peptidyl transferase activity comes from the 23S rRNA of the large subunit
Translation rate: approximately 6 amino acids per second in eukaryotes; approximately 20 per second in bacteria
Initiation
A special initiator tRNA, charged with methionine, is loaded into the P site of the small ribosomal subunit along with translation initiation factors
The small subunit (with initiator tRNA) binds to the 5' end of the mRNA, marked by the 5' cap in eukaryotes
It scans along the mRNA in the 5' to 3' direction until it encounters the first AUG codon
Initiation factors dissociate, the large ribosomal subunit joins, and translation begins
In prokaryotes: no 5' cap; instead, a ribosome-binding sequence (6 nucleotides upstream of AUG) allows ribosomes to bind directly to internal start codons. Prokaryotic mRNAs can be polycistronic (encoding multiple proteins from one mRNA)
Elongation (four-step cycle)
Step 1: a charged tRNA enters the vacant A site by base-pairing its anticodon with the exposed mRNA codon. Only a matching tRNA binds efficiently
Step 2: the polypeptide chain is uncoupled from the tRNA at the P site and joined by a peptide bond to the amino acid on the tRNA at the A site (catalysed by the large subunit)
Step 3: the large subunit shifts forward, moving the two tRNAs into the E and P sites
Step 4: the small subunit moves exactly three nucleotides along the mRNA, ejecting the spent tRNA from the E site and resetting the A site for the next charged tRNA
Termination
A stop codon (UAA, UAG, or UGA) enters the A site
No tRNA recognises it; instead, a release factor binds
The release factor causes addition of water to the polypeptide chain, freeing it from the tRNA
The ribosome releases the mRNA and dissociates into its two subunits, ready for another round
Multiple ribosomes can translate a single mRNA simultaneously, forming a polyribosome
In bacteria, ribosomes can begin translating an mRNA before transcription of that RNA is complete (coupled transcription-translation)
Tetracycline: blocks aminoacyl-tRNA binding to the A site
Streptomycin: prevents the transition from initiation to elongation, causes miscoding
Chloramphenicol: blocks the peptidyl transferase reaction
Cycloheximide: blocks the translocation step
Rifamycin: blocks initiation of transcription by inhibiting RNA polymerase
After translation, proteins must fold into the correct three-dimensional shape
Some fold spontaneously; most require chaperone proteins to guide folding and prevent aggregation
Completed polypeptides may also need covalent modifications (e.g. phosphorylation) or assembly with cofactors and protein partners
Controlled protein degradation regulates the amount of each protein in the cell
Proteases: enzymes that cut peptide bonds, degrading proteins to short peptides and then to individual amino acids
The proteasome: a large barrel-shaped complex in the cytosol and nucleus. ATP hydrolysis unfolds proteins and threads them into the central chamber where proteases chop them
Ubiquitin tags mark proteins for proteasomal degradation
Students often think DNA directly makes protein. It does not. DNA is transcribed into RNA, and RNA is translated into protein. DNA never leaves the nucleus in eukaryotes.
Students often confuse the template strand with the coding strand. The template strand is the one RNA polymerase reads (3' to 5'). The coding strand has the same sequence as the RNA product (with T instead of U) and is not read by the polymerase.
Students often assume that all 64 codons code for amino acids. Three of them (UAG, UAA, UGA) are stop codons and do not specify any amino acid.
Students often think that introns are "junk." While introns do not code for protein, their removal by splicing is essential for producing a functional mRNA, and alternative splicing can generate multiple different proteins from a single gene.
⚠️ Be able to trace the full path from gene to functional protein in a eukaryotic cell: transcription, capping, splicing, polyadenylation, nuclear export, translation, folding, modification.
⚠️ Know the differences between transcription and DNA replication (substrates, primer requirement, proofreading, product).
⚠️ Know the roles of sigma factor (bacteria) versus general transcription factors (eukaryotes) in transcription initiation.
⚠️ Understand the genetic code: 64 codons, redundancy, one start codon (AUG), three stop codons (UAG, UAA, UGA). Be able to use a codon table.
⚠️ Be able to describe the four-step translation elongation cycle and the roles of the A, P, and E sites.
⚠️ Understand that the ribosome is a ribozyme: rRNA catalyses peptide bond formation, not protein.
⚠️ Know how antibiotics exploit differences between prokaryotic and eukaryotic protein synthesis machinery.
True or False: RNA polymerase reads the coding strand to produce mRNA.
Fill in the blank: The three processing steps for eukaryotic mRNA are ________, ________, and ________.
True or False: The start codon AUG codes for methionine and sets the reading frame.
Fill in the blank: The enzyme that attaches the correct amino acid to its tRNA is called a(n) ________.
True or False: In eukaryotes, RNA polymerase II can initiate transcription without any accessory proteins.
Answers: 1. False (it reads the template strand). 2. 5' capping, RNA splicing, 3' polyadenylation. 3. True. 4. Aminoacyl-tRNA synthetase. 5. False (it requires general transcription factors).
Q: What is the central dogma of molecular biology, and what are the two main steps involved?
A: The central dogma states that genetic information flows from DNA to RNA to protein. The two main steps are transcription (DNA to RNA) and translation (RNA to protein).
Q: Describe three key differences between RNA polymerase and DNA polymerase.
A: RNA polymerase uses ribonucleoside triphosphates (not deoxyribonucleotides), can start a chain without a primer, and does not accurately proofread (error rate approximately 1 in 10,000).
Q: How does eukaryotic transcription initiation differ from prokaryotic transcription initiation?
A: Prokaryotes use a single RNA polymerase with a sigma factor that can recognise the promoter directly. Eukaryotes use three RNA polymerases (Pol II for protein-coding genes) and require a set of general transcription factors (TFIID, TFIIB, TFIIH, etc.) to assemble at the promoter before Pol II can begin.
Q: What are the roles of the 5' cap and poly-A tail on eukaryotic mRNA?
A: Both modifications increase mRNA stability and facilitate nuclear export. The 5' cap also marks the molecule as mRNA and helps recruit the ribosome during translation initiation.
Q: Explain the function of tRNA in translation.
A: tRNA acts as an adaptor molecule. Its anticodon base-pairs with a complementary codon on the mRNA, while its 3' end carries the corresponding amino acid. This allows the ribosome to match each mRNA codon to the correct amino acid during protein synthesis.
Q: What happens when a ribosome encounters a stop codon during translation?
A: No tRNA recognises the stop codon. Instead, a release factor binds to the A site, causing addition of water to the polypeptide chain. This frees the completed protein from the tRNA, and the ribosome dissociates into its two subunits.
Q: Why is the ribosome considered a ribozyme?
A: The ribosome is two-thirds RNA by weight, and its catalytic activity (peptide bond formation) is carried out by the rRNA component (specifically 23S rRNA in the large subunit), not by the ribosomal proteins. The proteins mainly stabilise the RNA core.
This material connects directly to Chapter 8 (gene regulation), which explains how cells control when and how much of each protein is produced by regulating transcription.
RNA processing and mRNA export link to Chapter 15 (intracellular compartments), which covers how proteins are sorted to the correct cellular locations after translation.
The antibiotic mechanisms described here connect to microbiology and pharmacology: drugs that target bacterial ribosomes work because prokaryotic and eukaryotic ribosomes are structurally different enough to allow selective inhibition.
central dogma, transcription, translation, mRNA, messenger RNA, tRNA, transfer RNA, rRNA, ribosomal RNA, RNA polymerase, RNA Pol II, promoter, TATA box, TFIID, TBP, TFIIH, sigma factor, general transcription factors, template strand, coding strand, codon, anticodon, start codon, AUG, stop codon, UAG, UAA, UGA, genetic code, reading frame, wobble base pairing, aminoacyl-tRNA synthetase, ribosome, A site, P site, E site, peptide bond, ribozyme, polyribosome, polysome, 5' cap, poly-A tail, RNA splicing, intron, exon, snRNP, spliceosome, lariat, pre-mRNA, mRNA processing, gene expression, protein synthesis, protein folding, chaperone, proteasome, ubiquitin, proteolysis, BIOL 101, cell biology, molecular biology