BioCentral IBBioCentral IB
IB Biology · Theme D · D1.2

Protein synthesis

A cell converts a DNA base sequence into a working protein in two connected stages: transcription copies a gene into mRNA, then translation reads that mRNA three bases at a time to build a chain of amino acids. Accuracy at every step — and a series of checks and modifications afterward — is what turns a genetic sequence into a functional molecule.
Guiding questions

How does a cell produce a sequence of amino acids from a sequence of DNA bases?

How is the reliability of protein synthesis ensured?

Part one

Transcription — copying a gene into RNA

D1.2.1 – D1.2.4
D1.2.1

Copying a gene into RNA

Transcription is the synthesis of RNA using a DNA template — the first step in reading a gene's information.
  • RNA polymerase binds the DNA template strand of a gene and moves along it, joining free ribonucleotides into a growing RNA strand.
  • Unlike DNA polymerase, RNA polymerase does not need a primer to start — it can begin synthesis on its own once bound to the DNA.
  • Only the section of DNA making up the gene being transcribed is copied; the rest of the DNA molecule stays untouched and double-stranded.
Why it mattersTranscription is the checkpoint that decides which genes actually get expressed as RNA, and eventually protein, at any given moment (D1.2.4).
A 3D molecular render of RNA polymerase, a large bumpy enzyme, sitting on a DNA double helix. Inside the enzyme the two DNA strands are separated into a transcription bubble, with the non-template (coding) strand above and the DNA template strand below, and one single new RNA strand leaves the enzyme.
D1.2.2

Base pairing builds the RNA strand

Hydrogen bonding and complementary base pairing between the DNA template and incoming ribonucleotides determine the exact sequence of the new RNA strand.
  • Each exposed template base pairs with its complementary ribonucleotide: C with G, T with A, and — the one exception from DNA replication — A with U.
  • RNA uses uracil instead of thymine, so adenine on the DNA template pairs with uracil, not thymine, in the growing RNA strand.
  • Because pairing is fixed and complementary, the RNA sequence produced is a predictable, faithful copy of the template's information, not a random assembly.
Why it mattersThis is the same base-pairing logic used in DNA replication (D1.1), applied here to build RNA instead of a second DNA strand.
Four DNA template bases A, T, G and C in a blue row, each paired by dashed hydrogen bonds with the RNA base directly below it: U, A, C and G in an orange row. The first pair is highlighted: adenine on the DNA pairs with uracil in RNA, not thymine.
D1.2.3

One template, many transcripts

A single DNA strand can be used as a template for transcription repeatedly, without changing its own base sequence.
  • Producing an RNA transcript doesn't consume or alter the template — the same gene can be transcribed again and again on demand.
  • In non-dividing somatic cells, this stability matters over an entire lifetime: the base sequence of these genes must be conserved for as long as that cell exists.
  • This is distinct from replication (D1.1), which produces a new, permanent DNA copy — transcription produces a disposable, temporary RNA copy instead.
Why it mattersA stable template means a cell can produce a protein's mRNA on demand, whenever and as often as that protein is needed.
One unchanged DNA double helix on the left with three arrows, one to each of three identical single-stranded RNA transcripts labelled 1st, 2nd and 3rd transcript, showing that the same template makes identical copies again and again.
D1.2.4

Switching genes on and off

Transcription is required for gene expression — a gene's information only becomes usable once it is transcribed into RNA.
  • Not every gene in a cell's genome is expressed at any given time; different cell types transcribe different subsets of genes despite sharing the same DNA.
  • Because transcription is the first stage of gene expression, it is a key stage at which a gene can be switched on or off.
  • A gene that is never transcribed produces no mRNA, and therefore no protein, regardless of what the DNA sequence itself specifies.
Why it mattersThis is why a liver cell and a neuron, with identical genomes, can look and function completely differently — it comes down to which genes get transcribed.
Inside a cell nucleus, two DNA double helices. On Gene 1 an RNA polymerase is transcribing and an RNA transcript leaves it (switched on). Gene 2 is bare with no enzyme and no RNA (switched off).
Part two

Translation — building the polypeptide

D1.2.5 – D1.2.11
D1.2.5

From RNA sequence to amino acids

Translation is the synthesis of polypeptides from mRNA — the base sequence of mRNA is translated into the amino acid sequence of a polypeptide.
  • Translation takes place at ribosomes, where the mRNA's base sequence is read and converted into a specific sequence of amino acids.
  • This conversion moves information from the language of nucleotide bases to the entirely different language of amino acids — the genetic code (D1.2.8) is what makes that translation possible.
  • The resulting chain of amino acids, the polypeptide, then folds and is often modified further before it can function (D1.2.18).
Why it mattersTranslation is where genetic information actually becomes a working molecule — proteins do the cell's structural and enzymatic work, not DNA or RNA directly.
A 3D molecular render of a ribosome, a large subunit on top of a smaller one, with an mRNA strand passing between them, one tRNA delivering an amino acid, and a growing polypeptide chain of amino acid beads leaving the top.
D1.2.6

Three players, one ribosome

mRNA, ribosomes and tRNA each play a distinct role in translation.
  • mRNA carries the coded instructions and binds to the small subunit of the ribosome, which reads it codon by codon.
  • tRNA molecules carry specific amino acids and match them to the mRNA sequence through base pairing; two tRNAs can bind simultaneously to the ribosome's large subunit.
  • The ribosome itself provides the structure that holds mRNA and tRNA in the right positions and catalyses the peptide bonds linking amino acids together.
Why it mattersTranslation only works because all three components are held together in one place — the ribosome is essentially the ordering and assembly hub.
A 3D molecular render of a ribosome. The large subunit on top is clearly bigger than the small subunit below. The mRNA (5' end left, 3' end right) runs along the small subunit, and two tRNA molecules sit side by side under the large subunit: the tRNA in the P site on the left and the tRNA in the A site on the right.
D1.2.7

Codon meets anticodon

Complementary base pairing between tRNA and mRNA is what matches each amino acid to the correct position in the growing polypeptide.
  • A codon is a triplet of three consecutive bases on the mRNA strand; an anticodon is the complementary triplet of bases on a tRNA molecule.
  • A tRNA can only bind an mRNA codon if its anticodon is complementary to that codon — this pairing is what selects the correct amino acid for each position.
  • Each tRNA molecule carries one specific amino acid, attached at the end opposite its anticodon loop, ready to be added to the chain.
Why it mattersCodon-anticodon pairing is the direct physical link between a gene's base sequence and a protein's amino acid sequence.
A tRNA molecule with an amino acid attached at its top end. Its anticodon U A C sits above the mRNA codon A U G, joined by three dashed lines showing complementary base pairing.
D1.2.9

Reading the code from a table

The genetic code is provided in the data booklet as the complete table below, letting the amino acid sequence coded by any mRNA strand be deduced directly.
First base Second base Third base
U C A G
U UUU phenylalanine UCU serine UAU tyrosine UGU cysteine U
UUC phenylalanine UCC serine UAC tyrosine UGC cysteine C
UUA leucine UCA serine UAA stop ** UGA stop ** A
UUG leucine UCG serine UAG stop ** UGG tryptophan G
C CUU leucine CCU proline CAU histidine CGU arginine U
CUC leucine CCC proline CAC histidine CGC arginine C
CUA leucine CCA proline CAA glutamine CGA arginine A
CUG leucine CCG proline CAG glutamine CGG arginine G
A AUU isoleucine ACU threonine AAU asparagine AGU serine U
AUC isoleucine ACC threonine AAC asparagine AGC serine C
AUA isoleucine ACA threonine AAA lysine AGA arginine A
AUG methionine* ACG threonine AAG lysine AGG arginine G
G GUU valine GCU alanine GAU aspartate GGU glycine U
GUC valine GCC alanine GAC aspartate GGC glycine C
GUA valine GCA alanine GAA glutamate GGA glycine A
GUG valine GCG alanine GAG glutamate GGG glycine G
*Note: AUG is an initiator codon and also codes for the amino acid methionine.** Note: UAA, UAG and UGA are terminator codons.
Why it mattersThis exact table is in the real data booklet — reading it fluently turns a string of mRNA bases into a predictable amino acid sequence, a core exam skill for this topic.
D1.2.8

Reading three bases at a time

The genetic code specifies each amino acid using a triplet of three consecutive mRNA bases, and has two further defining features: degeneracy and universality.
  • A triplet code is required because there are 20 amino acids but only 4 bases: neither single bases (4) nor pairs of bases (4²=16) give enough combinations, but triplets (4³=64) comfortably cover all 20.
  • Degeneracy means most amino acids are specified by more than one codon — look back at the table you just read: GGU, GGC, GGA and GGG all code for the same amino acid, glycine.
  • Universality means the same codon specifies the same amino acid in essentially all organisms, from bacteria to plants to humans.
Why it mattersUniversality is strong evidence that all life shares a common evolutionary origin, and is exactly what makes tools like genetic engineering possible across species.
Degeneracy: four codons GGU, GGC, GGA and GGG all lead to the amino acid glycine. Universality: photographs of a bacterium, a plant and a human all lead to the codon AUG and the amino acid methionine.
D1.2.10

The ribosome moves, the chain grows

The ribosome moves stepwise along the mRNA, one codon at a time, linking each new amino acid to the growing polypeptide chain by a peptide bond.
  • At each step, a tRNA carrying the next amino acid binds to the codon currently in the ribosome, and a peptide bond forms between it and the last amino acid in the chain.
  • The ribosome then moves exactly one codon further along the mRNA, and the process repeats — elongating the polypeptide one amino acid at a time.
  • This stepwise, codon-by-codon movement is what keeps amino acids added in precisely the order specified by the mRNA sequence.
Why it mattersThis focuses on elongation specifically — the repeating middle stage of translation, distinct from how it starts (D1.2.17, HL) or ends.
A ribosome on an mRNA strand, with the 5' end on the left and the 3' end on the right and an arrow showing that the ribosome moves toward the 3' end, one codon at a time. Two tRNAs sit in the ribosome and a chain of amino acid beads leaves the top, the newest amino acid joined to the chain by a peptide bond.
D1.2.11

One base changed, one protein altered

A mutation that changes even a single DNA base can change the resulting protein's structure by altering which amino acid gets placed in the polypeptide.
  • A point mutation substitutes one base for another; if it falls within a codon, it can change which amino acid that codon specifies.
  • For example, changing the middle base of the codon GAG (glutamate) to GUG changes the amino acid to valine — a single substitution, one changed amino acid.
  • A different amino acid can change how the polypeptide folds, which can disrupt the protein's shape and therefore its function — as in sickle cell haemoglobin.
Why it mattersThis links DNA-level change directly to a functional consequence — a single base substitution is enough to alter an entire protein's behaviour.
Two rows. Normal: the codon GAG gives glutamate, and normal haemoglobin gives round, flexible red blood cells (electron micrograph style image). Mutated: the middle base changes so the codon is GUG, giving valine, and sickle cell haemoglobin gives crescent-shaped sickled red blood cells.
Quick check

A biologist finds that the codons CGA and CGC both code for the same amino acid, arginine — and that this is true in humans, bacteria, and plants alike. Which two features of the genetic code does this observation illustrate?

Degeneracy (more than one codon per amino acid) and universality (the same code across species)
The triplet nature of the code and complementary base pairing
Alternative splicing and post-transcriptional modification
Semi-conservative replication and proofreading
Correct answer: degeneracy and universality. CGA and CGC coding for the same amino acid shows degeneracy — more than one codon per amino acid. The same pairing holding true across humans, bacteria and plants shows universality — the same codon means the same amino acid in essentially all organisms.
Part three

HL — directionality, promoters and non-coding DNA

D1.2.12 – D1.2.14
D1.2.12 · HL

Both processes run 5' to 3'

Transcription and translation both proceed with a fixed direction: RNA polymerase reads the DNA template 3' to 5' while building a new RNA strand 5' to 3', and the ribosome reads mRNA 5' to 3'.
  • RNA polymerase moves along the DNA template strand from its 3' end toward its 5' end, which means the new mRNA strand it builds necessarily grows in the 5' to 3' direction.
  • During translation, the ribosome reads the finished mRNA strand starting near its 5' end and moves toward its 3' end, one codon at a time.
  • Transcription follows the same rule as DNA replication (D1.1): a new nucleic acid strand is only ever extended 5' to 3'. Translation then reads that mRNA in the same 5' to 3' direction.
Why it mattersThis consistent directionality is what makes the mRNA sequence read the same way, start to finish, by every ribosome that translates it.
Left, transcription: RNA polymerase on the DNA template strand, which runs 3' on the left to 5' on the right, with arrows showing the enzyme moving rightward and the new mRNA growing 5' to 3'. Right, translation: a ribosome on mRNA with the 5' end on the left and the 3' end on the right, moving rightward, 5' to 3'.
D1.2.13 · HL

Starting transcription at the promoter

Transcription begins when transcription factors bind to the promoter, a specific DNA sequence found near the start of a gene.
  • Transcription factors are proteins that recognize and bind to the promoter sequence before RNA polymerase itself becomes involved.
  • Once bound, these transcription factors allow RNA polymerase to bind the promoter region too, positioning it correctly to begin transcription just downstream.
  • Naming specific transcription factors isn't required at this level — the key idea is that binding at the promoter is what initiates transcription at the right place.
Why it mattersBecause the promoter also determines whether a gene can be transcribed at all, it is a key control point in switching gene expression on (D1.2.4).
A DNA double helix with two transcription factor proteins bound at the promoter, an RNA polymerase enzyme sitting just beside them, and an arrow showing that transcription then moves along the DNA to the right.
D1.2.14 · HL

DNA that doesn't code for protein

Large parts of a eukaryotic genome are non-coding sequences that do not code for polypeptides, even though they still have real biological roles.
  • Regulators of gene expression control when and how strongly nearby genes are transcribed, without coding for a polypeptide themselves.
  • Introns are non-coding sequences within a gene that get removed during post-transcriptional modification (D1.2.15); telomeres are repetitive non-coding sequences that protect the ends of chromosomes.
  • Genes for rRNA and tRNA are also non-coding for polypeptides — they're transcribed into functional RNA molecules, not translated into protein at all.
Why it mattersMost of a eukaryotic genome is non-coding — protein-coding genes make up only a small fraction of total DNA.
One stretch of eukaryotic DNA drawn as five coloured regions: a protein-coding gene, a regulator of gene expression, an intron, a gene for rRNA or tRNA, and a telomere at the end. Only the first region codes for a polypeptide; the other four do not.
Part four

HL — processing mRNA and finishing the protein

D1.2.15 – D1.2.19
D1.2.15 · HL

Editing the transcript before it's used

In eukaryotic cells, the initial RNA transcript is modified after transcription before it becomes a mature, usable mRNA molecule.
  • Introns are removed from the transcript, and the remaining exons are spliced together to form the continuous coding sequence of the mature mRNA.
  • A 5' cap is added to the transcript's 5' end, and a 3' poly(A) tail — a long run of adenine nucleotides — is added to its 3' end.
  • Both additions help stabilize the mRNA transcript, protecting it from degradation before and during translation.
Why it mattersOnly fully processed, mature mRNA leaves the nucleus for translation — unprocessed transcripts with introns still attached are not translated.
A pre-mRNA with three exons and two introns between them. Splicing removes the introns and joins the exons into a mature mRNA, which also has a 5' cap at its start and a 3' poly(A) tail at its end.
D1.2.16 · HL

One gene, several proteins

Alternative splicing joins different combinations of exons from the same pre-mRNA, allowing a single gene to code for more than one polypeptide variant.
  • Because which exons are included or skipped can vary between splicing events, the same starting transcript can be spliced into more than one distinct mature mRNA.
  • Each different exon combination translates into a different polypeptide sequence — a different protein variant from a single gene.
  • Specific real examples aren't required at this level; the key idea is that splicing choice, not just gene number, increases the range of proteins a genome can produce.
Why it mattersAlternative splicing is a major reason the human genome codes for far more distinct proteins than it has genes.
One pre-mRNA with four coloured exons numbered 1 to 4. Splicing gives mRNA variant A (exons 1, 2, 4; exon 3 skipped) or mRNA variant B (exons 1, 3, 4; exon 2 skipped), which are translated into two different protein variants.
D1.2.17 · HL

Assembling the ribosome to start

Translation begins with a defined initiation sequence: the small ribosomal subunit attaches to the mRNA's 5' terminal, then moves along it until it reaches the start codon.
  • At the start codon, an initiator tRNA binds, followed by a second tRNA binding the next codon; only then does the large ribosomal subunit attach, completing the ribosome.
  • The assembled ribosome has three binding sites for tRNA — the A, P and E sites — that operate during the elongation stage that follows initiation.
  • During elongation, a new tRNA enters at the A site, the growing chain transfers there from the P site, and the empty tRNA leaves via the E site.
Why it mattersStarting at the correct codon keeps the whole downstream reading frame, and therefore the whole protein sequence, correct.
Three panels. 1: the small ribosomal subunit sits on the mRNA and moves from the 5' end toward the AUG start codon. 2: an initiator tRNA and a second tRNA are bound while the large subunit approaches. 3: the large subunit has joined, and the ribosome shows the E, P and A tRNA sites.
D1.2.18 · HL

A protein isn't finished at translation

Many polypeptides must be modified after translation before they can function — pre-proinsulin's two-stage modification into insulin is a defined example.
  • Pre-proinsulin, the initial translation product, has its signal peptide removed to produce proinsulin — a first modification stage.
  • Proinsulin then has its central C-peptide section removed, leaving two separate chains, the A and B chains, still linked together by disulfide bonds.
  • This second modification stage produces mature, functional insulin — the molecule cannot lower blood glucose in either of its earlier, unmodified forms.
Why it mattersPost-translational modification means the amino acid sequence a ribosome builds is often not the final functional molecule — extra processing steps are frequently still required.
Three stages drawn as chains of beads. Pre-proinsulin has a signal peptide, a B chain, a C-peptide and an A chain. Removing the signal peptide gives proinsulin, folded with the C-peptide as a loop and two disulfide bonds between the chains. Removing the C-peptide leaves insulin: separate A and B chains held together by disulfide bonds.
D1.2.19 · HL

Breaking proteins down to build new ones

Proteasomes break down proteins that are damaged or no longer needed, recycling their amino acids for use in synthesizing new proteins.
  • Sustaining a functional proteome requires constant protein breakdown alongside constant protein synthesis — old or faulty proteins are continually cleared out, not left to accumulate.
  • The amino acids released by proteasome breakdown re-enter the cell's free amino acid pool, available to be used again in translating new polypeptides.
  • This constant turnover lets a cell adjust its protein content in response to changing conditions, rather than being stuck with whatever it made previously.
Why it mattersRecycling amino acids this way is far more efficient than relying solely on new amino acids from digestion or fresh synthesis.
A tangled damaged protein heads toward a barrel-shaped proteasome, which breaks it into free amino acid beads. An arrow shows the amino acids being reused to build a new protein.
Quick check · HL

A mutation destroys a gene's promoter but leaves the rest of the gene, including its start codon, completely unchanged. What is the most likely direct effect?

Transcription factors and RNA polymerase can no longer bind, so the gene is not transcribed at all
The mRNA will be translated starting at the wrong codon
The protein will fold normally but lack a signal peptide
Alternative splicing will remove the promoter's introns
Correct answer: the gene isn't transcribed. Transcription factors bind the promoter to allow RNA polymerase to bind and begin transcription; without a functional promoter, that initiation step can't happen, so no mRNA — and therefore no protein — can be produced from that gene, regardless of the coding sequence itself being intact.

Key vocabulary

Worth being able to define in a single sentence each

Codon
A triplet of three consecutive bases on mRNA that specifies a particular amino acid or a start/stop signal.
Anticodon
The complementary triplet of bases on a tRNA molecule that pairs with a codon on mRNA.
Degeneracy
The genetic code's use of more than one codon to specify most amino acids.
Universality
The genetic code's property that the same codon specifies the same amino acid across essentially all organisms.
Point mutation
A change to a single base in a DNA sequence, which can alter the amino acid a codon specifies.
PromoterHL
A DNA sequence near the start of a gene where transcription factors and RNA polymerase bind to begin transcription.
IntronHL
A non-coding sequence within a gene that is removed from the transcript during post-transcriptional modification.
ProteasomeHL
A cell structure that breaks down damaged or unneeded proteins, releasing amino acids for reuse.
C-peptideHL
The central section removed from proinsulin, leaving the linked A and B chains of mature insulin.

Where this shows up again

A1.2 / D1.1
Transcription uses DNA (A1.2, D1.1) as a template. Explain why only one DNA strand (the template strand) is transcribed for a given gene, and how DNA's double-stranded structure protects the template.
B1.2
Ribosomes are composed of rRNA and protein, and have A, P, and E sites (B1.2). How does the quaternary structure of the ribosome enable its function in translation?
B2.2
Many proteins synthesized by ribosomes on the rough ER (B2.2) are destined for secretion or membrane insertion. Suggest how post-translational modification (D1.2.18) relates to a protein's eventual destination in the cell.
D1.3
Mutations (D1.3) that change a single base can have very different consequences depending on the genetic code's degeneracy (D1.2.8). Explain why a substitution mutation is more likely to be silent than an insertion or deletion of one base.

D1.2 Protein synthesis — one-page recap

Screenshot this slide to revise from

Transcription
  • RNA polymerase synthesizes RNA from a DNA template, pairing A on DNA with U on RNA (no primer needed).
  • The template is reusable and unchanged; transcription is the key step where a gene is switched on or off.
Translation basics
  • mRNA binds the small subunit; two tRNAs bind the large subunit at once.
  • Codon (mRNA triplet) pairs with the complementary anticodon (tRNA triplet) carrying one amino acid.
Genetic code & mutation
  • Triplet code (4³=64); degenerate (multiple codons/amino acid) and universal across organisms.
  • A point mutation can change one codon's amino acid and disrupt protein shape.
Ribosome mechanics
  • The ribosome moves stepwise, one codon at a time, toward the mRNA's 3' end.
  • Each step adds one amino acid via a new peptide bond to the growing chain.
HL · Directionality, promoters & non-coding DNA
  • Both processes run 5'→3'; transcription factors + RNA polymerase bind the promoter to start transcription.
  • Non-coding DNA includes regulators, introns, telomeres, and rRNA/tRNA genes.
HL · Processing & finishing the protein
  • Introns spliced out, 5' cap + polyA tail added; alternative splicing makes multiple proteins from one gene.
  • Initiation assembles the ribosome at the start codon (A/P/E sites); pre-proinsulin → insulin and proteasome recycling finish the story.

From a gene's sequence to a working protein

Two coupled processes, transcription and translation, turn a DNA base sequence into a specific chain of amino acids — and a series of checks, edits and modifications along the way is what keeps that chain accurate and turns it into a molecule that actually works.
D1.2 Protein synthesis · BioCentral IB
Use ↓ ↑ or click to navigate
01 / 31