CHAPTER 5
MOLECULAR BASIS OF INHERITANCE
THE DNA
DNA is a long polymer of deoxyribonucleotides.
The length of DNA is usually defined as number of nucleotides (or a pair of nucleotide referred to as base pairs)
present in it. This also is the characteristic of an organism.
1. A bacteriophage known as φ ×174 has 5386 nucleotides,
2. Bacteriophage lambda has 48502 base pairs (bp),
3. Escherichia coli has 4.6 × 106 bp,
4. Haploid content of human DNA is 3.3 × 109 bp.
STRUCTURE OF POLYNUCLEOTIDE CHAIN
The chemical structure of a polynucleotide chain (DNA or RNA). A nucleotide has
three components –
1. A nitrogenous base,
2. A pentose sugar (ribose in case of RNA, and deoxyribose for DNA),
3. A phosphate group.
There are two types of nitrogenous bases –
1. Purines (Adenine and Guanine),
2. Pyrimidines (Cytosine, Uracil and Thymine).
Cytosine is common for both DNA and RNA and Thymine is present in DNA. Uracil is present in RNA at the place
of Thymine.
BONDS:
A nitrogenous base is linked to the OH of 1' C pentose sugar through aN-glycosidic linkage to form a
nucleoside, such as adenosine or deoxyadenosine, guanosine or deoxyguanosine, cytidine or
deoxycytidine and uridine or deoxythymidine.
When a phosphate group is linked to OH of 5' C of a nucleoside through phosphoester linkage, a
corresponding nucleotide (or deoxynucleotide depending upon the type of sugar present) is formed.
Two nucleotides are linked through 3'-5' phosphodiester linkage to form dinucleotide.
More nucleotides can be joined in such a manner to form a
polynucleotide chain.
• A polymer thus formed has at one end a free phosphate
moiety at 5' -end of sugar, which is referred to as 5’-end of
polynucleotide chain.
• Similarly, at the other end of the polymer the sugar has a
free OH of 3'C group which is referred to as 3' -end of the
polynucleotide chain.
The backbone of a polynucleotide chain is formed due to sugar
and phosphates.
In RNA, every nucleotide residue has an additional –OH group present at 2' -position in the ribose.
Also, in RNA the uracil is found at the place of thymine (5-methyl uracil, another chemical name for
thymine).
DNA as an acidic substance present in nucleus was first identified by Friedrich Meischer in 1869.
He named it as ‘Nuclein’.
It was only in 1953 that James Watson and Francis Crick, based on the X-ray diffraction data produced by
Maurice Wilkins and Rosalind Franklin, proposed a very simple but famous Double Helix model for the
structure of DNA.
Erwin Chargaff that for a double stranded DNA, the ratios between Adenine and Thymine and Guanine and
Cytosine are constant and equals one.
THE SALIENT FEATURES OF THE DOUBLE-HELIX STRUCTURE OF DNA:
(i) It is made of two polynucleotide chains, where the backbone is constituted by sugar-phosphate, and the
bases project inside.
(ii)The two chains have anti-parallel polarity. It means, if one chain has the polarity 5'-- 3', the other has 3'-- 5'.
(iii)The bases in two strands are paired through hydrogen bond (H-bonds) forming base pairs (bp). Adenine
forms two hydrogen bonds with Thymine from opposite strand and vice-versa. Similarly, Guanine is bonded
with Cytosine with three Hbonds. As a result, always a purine comes opposite to a pyrimidine. This generates
approximately uniform distance between the two strands of the helix
(iv)The two chains are coiled in a right-handed fashion. The pitch of the helix is 3.4 nm (a nanometre is one
billionth of a metre, that is 10-9 m) and there are roughly 10 bp in each turn. Consequently, the distance
between a bp in a helix is approximately 0.34 nm.
(v)The plane of one base pair stacks over the other in double helix. This, in addition to H-bonds, confers
stability of the helical structure.
CENTRAL DOGMA:
Francis Crick proposed the Central dogma in molecular biology, which states that the genetic information
flows from DNA ------->RNA------- Protein. In some viruses the flow of information is in reverse direction,
that is, from RNA to DNA.
PACKAGING OF DNA HELIX
Taken the distance between two consecutive base pairs as 0.34 nm (0.34×10–9 m), if the length of DNA
double helix in a typical mammalian cell is calculated (simply by multiplying the total number of bp with
distance between two consecutive bp, that is, 6.6 × 109 bp × 0.34 × 10-9m/bp), it comes out to be
approximately 2.2 metres.
A length that is far greater than the dimension of a typical nucleus (approximately 10–6 m). In prokaryotes,
such as, E. coli, though they do not have a defined nucleus, the DNA is not scattered throughout the cell.
Packaging of DNA in Prokaryotes:
In prokaryotes, such as, E. coli, though they do not have a defined nucleus, the DNA is not scattered
throughout the cell.
DNA (being negatively charged) is held with some proteins (that have positive charges) in a region
termed as ‘nucleoid’.
The DNA in nucleoid is organised in large loops held by proteins.
Packaging of DNA in Eukaryotes:
In eukaryotes there is a set of positively charged, basic proteins called histones. A protein acquires charge
depending upon the abundance of amino acids residues with charged side chains.
Histones are rich in the basic amino acid residues lysine and
arginine. Both the amino acid residues carry positive charges in
their side chains.
Histones are organised to form a unit of eight molecules called
histone octamer.
The negatively charged DNA is wrapped around the positively
charged histone octamer to form a structure called nucleosome .
A typical nucleosome contains 200 bp of DNA helix.
Nucleosomes constitute the repeating unit of a threadlike,
stained (coloured) bodies in nucleus called chromatin.
The nucleosomes in chromatin are seen as ‘beads-onstring’ structure when viewed under electron
microscope (EM) .
The beads-on-string structure in chromatin is packaged to form chromatin fibers that are further
coiled and condensed at metaphase stage of cell division to form chromosomes.
The packaging of chromatin at higher level requires additional set of proteins that collectively are
referred to as Non-histone Chromosomal (NHC) proteins.
In a typical nucleus, some region of chromatin are loosely packed (and stains light) and are referred to
as euchromatin. The chromatin that is more densely packed and stains dark are called as
Heterochromatin.
Euchromatin is said to be transcriptionally active chromatin, whereas heterochromatin is inactive.
THE SEARCH FOR GENETIC MATERIAL .
By 1926, the quest to determine the mechanism for genetic inheritance had reached the molecular level.
Previous discoveries by Gregor Mendel, Walter Sutton, Thomas Hunt Morgan and numerous other scientists
had narrowed the search to the chromosomes located in the nucleus of most cells.
GRIFFITH EXPERIMENT WHICH LEAD TO THE DISCOVERY OF TRANSFORMING PRINCIPLE IN BACTERIA.
In 1928, Frederick Griffith, did experiments with Streptococcus pneumoniae which causes
pneumonia, and witnessed a miraculous transformation in the bacteria.
During the experiment, a living
organism (bacteria) had changed
in physical form.
When Streptococcus
pneumoniae are grown on a culture
plate, some produce smooth shiny
colonies (S type) due to mucous
(polysaccharide) coat ;while others
produce rough colonies (R type)
without this coat.
Mice infected with the S
strain (virulent) die from pneumonia
infection but mice infected with the R
strain do not develop pneumonia.
Griffith was able to kill
bacteria by heating them. He observed that heat-killed S strain bacteria injected into mice did not kill
them.
When he injected a mixture of heat-killed S and live R bacteria, the mice died. Moreover, he recovered
living S bacteria from the dead mice.
Conclusion: Based on the observation, Griffith concluded that R strain bacteria had been transformed by S
strain bacteria. The R strain inherited some ‘transforming principle’ from the heat-killed S strain bacteria
which made them virulent. And he assumed this transforming principle as genetic material.
BIOCHEMICAL CHARACTERISATION OF TRANSFORMING PRINCIPLE
Prior to the work of Oswald Avery, Colin MacLeod and Maclyn McCarty (1933-44), the genetic material was
thought to be a protein.
1. They worked to determine the biochemical nature of ‘transforming principle’ in Griffith's experiment.
They purified biochemicals (proteins, DNA, RNA, etc.) from the heat-killed S cells to see which ones
could transform live R cells into S cells.
2. They discovered that DNA alone from S bacteria caused R bacteria to become transformed.
3. They also discovered that protein-digesting enzymes (proteases) and RNA-digesting enzymes (RNases)
did not affect transformation, so the transforming substance was not a protein or RNA.
4. Digestion with DNase did inhibit transformation, suggesting that the DNA caused the transformation.
They concluded that DNA is the hereditary material, but not all biologists were convinced.
THE GENETIC MATERIAL IS DNA –HERSHEY CHASE EXPERIMENT:
The unequivocal proof that DNA is the genetic material came from the experiments of Alfred Hershey and
Martha Chase (1952).
Alfred Hershey and Martha Chase (1952) gave unequivocal proof that DNA is the genetic material.
In their experiments, bacteriophages (viruses that infect bacteria) were used.
They grew some viruses on a medium that contained radioactive phosphorus and some others on
radioactive sulphur containing medium.
Viruses grown in the presence of radioactive phosphorus contained radioactive DNA but not
radioactive protein because DNA contains phosphorus but protein does not. In the same way, viruses
grown on radioactive sulphur contained radioactive protein, but not radioactive DNA because DNA
does not contain sulphur.
Infection: Radioactive phages were
allowed to attach to E. coli bacteria.
Blending: Then, as the infection
proceeded, the viral coats were removed from
the bacteria by agitating them in a blender.
Centrifugation:The
virus particles were separated from the bacteria
by spinning them in a centrifuge.
Bacteria which was infected with viruses
that had radioactive DNA were radioactive,
indicating that DNA was the material that passed
from the virus to the bacteria.
Bacteria that were infected with viruses that had radioactive proteins were not radioactive. This
indicates that proteins did not enter the bacteria from the viruses.
DNA is therefore the genetic material that is passed from virus to bacteria.
PROPERTIES OF GENETIC MATERIAL (DNA VERSUS RNA)
Criteria to be fulfilled by a molecule to be the genetic material
(i)It should be able to generate its replica (Replication).
(ii)It should be stable chemically and structurally.
(iii) It should provide the scope for slow changes (mutation) that are required for evolution.
(iv) It should be able to express itself in the form of 'Mendelian Characters’.
RNA WORLD
• RNA was the first genetic material. There is now enough evidence to suggest that essential life
processes (such as metabolism, translation, splicing, etc.), evolved around RNA.
• RNA used to act as a genetic material as well as a catalyst (there are some important biochemical
reactions in living systems that are catalysed by RNA catalysts and not by protein enzymes).
• But, RNA being a catalyst was reactive and hence unstable. Therefore, DNA has evolved from RNA
with chemical modifications that make it more stable.
• DNA being double stranded and having complementary strand further resists changes by evolving a
process of repair.
REPLICATION
1. While proposing the double helical structure for DNA, Watson and Crick had immediately
proposed a scheme for replication of DNA.
2. To quote their original statement that is as follows: ‘‘It has not escaped our notice that the
specific pairing we have postulated immediately suggests a possible copying mechanism for
the genetic material’’ (Watson and Crick, 1953).
3. The scheme suggested that the two strands would separate and act as a template for the
synthesis of new complementary strands.
After the completion of replication, each DNA molecule would have one parental and one newly
synthesised strand. This scheme was termed as semiconservative DNA replication .
MESELSON AND STAHL EXPERIMENT TO PROVE THAT REPLICATION IS SEMICONSERVATIVE
It is now proven that DNA replicates semiconservatively. It was shown first in
Escherichia coli and subsequently in higher organisms, such as plants and human cells.
Matthew Meselson and Franklin Stahl performed the following experiment in 1958:
(i) They grew E. coli in a medium containing 15NH4Cl (15N is the heavy isotope of nitrogen) as the only
nitrogen source for many generations. The result was that 15N was incorporated into newly
synthesised DNA (as well as other nitrogen containing compounds).
(ii) This heavy DNA molecule could be distinguished from the normal DNA by centrifugation in a cesium
chloride (CsCl) density gradient (Please note that 15N
is not a radioactive isotope, and it can be separated from 14N only based on densities).
(iii) (ii) Then they transferred the cells into a medium with normal 14NH4Cl and took samples at various
definite time intervals as the cells multiplied, and extracted the DNA that remained as double-
stranded helices.
(iv) The various samples were separated independently on CsCl gradients to measure the densities of
DNA (Figure 6.7).
(v) Thus, the DNA that was extracted from the culture one generation after the transfer from 15N to
14N medium [that is after 20 minutes; E. coli divides in 20 minutes] had a hybrid or intermediate
density.
(vi) DNA extracted from the culture after another generation [that is after 40 minutes, II generation]
was composed of equal amounts of this hybrid DNA and of ‘light’ DNA.
(vii) Very similar experiments involving use of radioactive thymidine to detect distribution of newly
synthesised DNA in the chromosomes was performed on Viciafaba (faba beans) by Taylor and
colleagues in 1958. The experiments proved that the DNA in chromosomes also replicate
semiconservatively.
THE MACHINERY AND THE ENZYMES
•
In living cells, such as E. coli, the process of replication requires a set of catalysts (enzymes).
a)DNA-dependent DNA polymerase :
• It uses a DNA template to catalyse the polymerisation of deoxynucleotides. These enzymes are
highly efficient enzymes as they have to catalysepolymerisation of a large number of nucleotides
in a very short time.
• E. coli that has only 4.6 ×106bp (compare it with human whose diploid content is 6.6 × 109bp),
completes the process of replication within 18 minutes; that means the average rate of
polymerisation has to be approximately 2000 bp per second.
• Not only do these polymerases have to be fast, but they also have to catalyse the reaction with
high degree of accuracy.
• Any mistake during replication would result into mutations.
• Furthermore, energetically replication is a very expensive process.
b)Deoxyribonucleoside triphosphates : Serve dual purposes. In addition to acting as substrates, they provide
energy for polymerisation reaction (the two terminal phosphates in a deoxynucleoside triphosphates are
high-energy phosphates, same as in case of ATP).
c) DNA ligase:The discontinuously synthesised fragments are later joined by the enzyme DNA ligase
DNA REPLICATION PROCESS:
For long DNA molecules, since the two strands of DNA cannot be separated in its entire
length (due to very high energy requirement), the replication occur within a small opening
of the DNA helix, referred to as replication fork.
The DNA-dependent DNA polymerases uses a DNA template to catalyse the polymerisation
of deoxynucleotides. The DNA-dependent DNA polymerases catalysepolymerisation only in
one direction, that is 5'-- 3'. This creates some additional complications at the replicating
fork.
Consequently, on one strand (the template with polarity 3'-- 5'), the replication is
continuous, while on the other (the template with polarity 5'-- 3'), it is discontinuous.
The discontinuously synthesised fragments are later joined by the enzyme DNA ligase.
The DNA polymerases on their own cannot initiate the process of replication. Also the
replication does not initiate randomly at any place in DNA. There is a definite
region in E. coli DNA where the replication originates. Such regions are termed as origin of replication.
It is because of the requirement of the origin of replication that a
piece of DNA if needed to be propagated during recombinant DNA
procedures, requires a [Link] vectors provide the origin of
replication.
In eukaryotes, the replication of DNA takes place at S-phase of
the cell-cycle .The replication of DNA and cell division cycle should be
highly coordinated.A failure in cell
division after DNA replication results into polyploidy(a chromosomal
anomaly).
TRANSCRIPTION
• The process of copying genetic information from one strand of the DNA into RNA is termed as
transcription.
• Here also, the principle of complementarity governs the process of transcription, except the
adenosine complements now forms base pair with uracil instead of thymine.
• However, unlike in the process of replication, which once set in, the total DNA of an organism gets
duplicated, in transcription only a segment of DNA and only one of the strands is copied into RNA.
WHY BOTH THE STRANDS ARE NOT COPIED DURING TRANSCRIPTION ?
1. First, if both strands act as a template, they would code for RNA molecule with different
sequences (Remember complementarity does not mean identical), and in turn, if they code for
proteins, the sequence of amino acids in the proteins would be different. Hence, one segment of
the DNA would be coding for two different proteins, and this would complicate the genetic
information transfer machinery.
2. Second, the two RNA molecules if produced simultaneously would be complementary to each
other, hence would form a double stranded RNA. This would prevent RNA from being translated
into protein and the exercise of transcription would become a futile one.
TRANSCRIPTION UNIT:
A transcription unit in DNA is defined primarily by the three regions in the DNA:
(i) A Promoter ii) The Structural gene iii) A Terminator
• There is a convention in defining the two strands of the DNA in the structural gene of a transcription
unit.
• Since the two strands have opposite polarity and the DNA-dependent RNA polymerase also catalyse
the polymerisation in only one direction, that is, 5'→3', the strand that has the polarity 3'→5' acts as
a template, and is also referred to as template strand.
• The other strand which has the polarity (5'→3') and the sequence same as RNA (except thymine at the
place of uracil), is displaced during transcription. This strand (which does not code for anything) is
referred to as coding strand.
• The promoter and terminator flank the structural gene in a transcription unit.
• The promoter is said to be located towards 5' -end (upstream) of the structural gene (the reference is
made with respect to the polarity of coding strand). It is a DNA sequence that provides binding site for
RNA polymerase.
• The presence of a promoter in a transcription unit also defines the template and coding strands. By
switching its position with terminator, the definition of coding and template strands could be reversed.
• The terminator is located towards 3' -end (downstream) of the coding strand and it usually defines
the end of the process of transcription.
• There are additional regulatory sequences that may be present further upstream or downstream to
the promoter.
GENE: A gene is defined as the functional unit of inheritance.
The DNA sequence coding for tRNA or rRNA molecule also define a gene.
• However by defining a cistron as a segment of DNA coding for a polypeptide, the structural gene in a
transcription unit could be said as monocistronic (mostly in eukaryotes) or polycistronic (mostly in
bacteria or prokaryotes).
SPLIT GENE:
In eukaryotes, the monocistronic structural genes have interrupted coding sequences – the genes in
eukaryotes are split.
The coding sequences or expressed sequences are defined as exons. Exons are said to be those
sequence that appear in mature or processed RNA.
The exons are interrupted by introns. Introns or intervening sequences do not appear in mature or
processed RNA.
TYPES OF RNA AND THE PROCESS OF TRANSCRIPTION IN BACTERIA
In bacteria, there are three major types of RNAs:
i)mRNA (messenger RNA): The mRNA provides the template for protein synthesis.
ii) tRNA (transfer RNA):tRNA brings aminoacids to the site of protein synthesis and reads the genetic
code
iii) rRNA (ribosomal RNA): They play structural and catalytic role during translation.
.
PROCESS OF TRANSCRIPTION IN BACTERIA OR PROKARYOTES:
• There is single DNA-dependent RNA polymerase that catalyses transcription of all types of RNA in
bacteria.
• Initiation: RNA polymerase binds to promoter and initiates transcription.
• Elongation: It uses nucleoside triphosphates as substrate and polymerises in a template depended
fashion following the rule of complementarity. It somehow also facilitates opening of the helix and
continues elongation. Only a short stretch of RNA remains bound to the enzyme.
• Termination:Once the polymerases reaches the terminator region, the nascent RNA falls off, so also
the RNA polymerase. This results in termination of transcription.
• The RNA polymerase is only
capable of catalysing the process of
elongation. It associates transiently
with initiation-factor (σ) and
termination-factor (ρ) to initiate and
terminate the transcription,
respectively.
• Association with these factors
alter the specificity of the RNA
polymerase to either initiate or
terminate.
• In bacteria, since the mRNA
does not require any processing to
become active, and also since
transcription and translation take
place in the same compartment
(there is no separation of cytosol and
nucleus in bacteria), many times the translation can begin much before the mRNA is fully transcribed.
Consequently, the transcription and translation can be coupled in bacteria.
TRNSCRIPTION IN EUKARYOTES
In eukaryotes, there are two additional complexities –
i)There are at least three RNA polymerases
in the nucleus. There is a clear cut division
of labour.
• The RNA polymerase I transcribes rRNAs
(28S, 18S, and 5.8S),
• whereas the RNA polymerase III is
responsible for transcription of tRNA,
5srRNA, and snRNAs (small nuclear RNAs).
• The RNA polymerase II transcribes
precursor of mRNA, the heterogeneous
nuclear RNA (hnRNA).
ii)The second complexity is that the primary transcripts contain both the exons and the introns and are non-
functional.
• Hence, it is subjected to a process called splicing where the introns are removed and exons are joined
in a defined order.
• hnRNA undergoes additional processing called as capping and tailing.
• In capping an unusual nucleotide (methyl guanosine triphosphate) is added to the 5'end of hnRNA.
• In tailing, adenylate residues (200-300) are added at 3'-end in a template independent manner.
• It is the fully processed hnRNA, now called mRNA, that is transported out of the nucleus for
translation
. GENETIC CODE
1)George Gamow, a physicist, who argued that since there are only 4 bases and if they have to code for 20 amino
acids, the code should constitute a combination of bases. He suggested that in order to code for all the 20 amino
acids, the code should be made up of three nucleotides. This was a very bold proposition, because a permutation
combination of 43 (4 × 4 × 4) would generate 64 codons; generating many more codons than required. Providing
proof that the codon was a triplet, was a more daunting task.
2)The chemical method developed by Har Gobind Khorana was instrumental in synthesising RNA molecules with
defined combinations of bases (homopolymers and copolymers).
3)Marshall Nirenberg’s cell-free system for protein synthesis finally helped the code to be deciphered.
4)Severo Ochoa enzyme (polynucleotide phosphorylase) was also helpful in polymerising RNA with defined
sequences in a template independent manner (enzymatic synthesis of RNA).
THE SALIENT FEATURES OF GENETIC CODE:
1. The codon is triplet. 61 codons code for amino acids and 3 codons do not code for any amino acids,
hence they function as stop codons.
2. Some amino acids are coded by more than one codon, hence the code is degenerate.
3. The codon is read in mRNA in a contiguous fashion. There are no punctuations.
4. The code is nearly universal: for example, from bacteria to human UUU would code for
Phenylalanine (phe). Some exceptions to this rule have been found in mitochondrial codons, and in
some protozoans.
5. AUG has dual functions. It codes for Methionine (met) and it also act as initiator codon.
6. UAA, UAG, UGA are stop terminator codons.
MUTATIONS AND GENETIC CODE
Point mutation: is a change of single base pair in the gene.
Ex: Change of single base pair in the gene beta globin chain that results in the change of amino acid residue
glutamate to valine. It results into a diseased condition called as sickle cell anemia.
Frameshift mutation: Insertion or deletion of one or two bases changes the reading frame from the point of
insertion or deletion. So such mutations are called frameshift insertion or deletion mutations.
Insertion or deletion of three or its multiple bases insert or delete in one or multiple codon hence one or multiple
amino acids, and reading frame remains unaltered from that point onwards.
tRNA– THE ADAPTER MOLECULE
Francis Crick ,He postulated the presence of an adapter molecule that would on one hand read the
code and on other hand would bind to specific amino acids.
The tRNA, then called sRNA (soluble RNA), was known before the genetic code was postulated.
However, its role as an adapter molecule was assigned much later.
tRNA has an anticodon loop that has bases complementary to the code.
It also has an amino acid acceptor end to which it binds to amino acids.
tRNAs are specific for each amino acid .
For initiation, there is another specific tRNA that is referred to as initiator tRNA.
There are no tRNAs for stop codons.
The secondary structure of tRNA has been depicted that looks like a clover-leaf. In actual structure,
the tRNA is a compact molecule which looks like inverted
TRANSLATION:
Translation is the process of polymerisation of amino acids to form a polypeptide.
The order and sequence of amino acids are defined by the sequence of bases in the mRNA. The amino acids are
joined by a bond which is known as a peptide bond.
STEPS:
1)Charging of tRNA or aminoacylation of tRNA :
• Formation of a peptide bond requires energy. Therefore, in the first phase itself amino acids are
activated in the presence of ATP and linked to their cognate tRNA – a process commonly called as
charging of tRNA or aminoacylation of tRNA .
• If two such charged tRNAs are brought close enough, the formation of peptide bond between them
would be favoured energetically.
• The presence of a catalyst would enhance the rate of peptide bond formation. The cellular factory
responsible for synthesising proteins is the ribosome.
2)Initiation:
• When the small subunit encounters an mRNA, the process of translation of the mRNA to protein
begins.
• There are two sites in the large subunit, for subsequent amino acids to bind to and thus, be close
enough to each other for the formation of a peptide bond.
• A translational unit in mRNA is the sequence of RNA that is flanked by the start codon (AUG) and the
stop codon and codes for a polypeptide. An mRNA also has some additional sequences that are not
translated and are referred as untranslated regions (UTR).
• For initiation, the ribosome binds to the mRNA at the start codon (AUG) that is recognised only by the
initiator tRNA.
3)Elongation:
• The ribosome proceeds to the elongation phase of
protein [Link] this stage, complexes composed of an
amino acid linked to tRNA, sequentially bind to the appropriate
codon in mRNA by forming complementary base pairs with the
tRNA anticodon.
• The ribosome moves from codon to codon along the
mRNA. Amino acids are added one by one, translated into
Polypeptide sequences dictated by DNA and represented by
mRNA.
4)Termination:
• At the end, a release factor binds to the stop codon, terminating translation and releasing the
complete polypeptide from the ribosome.
REGULATION OF GENE EXPRESSION
. In eukaryotes, the regulation o gene could be exerted at different levels like
1. transcriptional level (formation of primary transcript),
2. processing level (regulation of splicing),
3. transport of mRNA from nucleus to the cytoplasm,
4. translational level
In a transcription unit, the activity of RNA polymerase at a given promoter is in turn regulated by
interaction with accessory proteins, which affect its ability to recognise start sites. These regulatory
proteins can act both positively (activators) and negatively (repressors).
The accessibility of promoter regions of prokaryotic DNA is in many cases regulated by the
interaction of proteins with sequences termed operators.
The operator region is adjacent to the promoter elements in most operons and in most cases the
sequences of the operator bind a repressor protein.
Each operon has its specific operator and specific repressor.
THE LAC OPERON
The elucidation of the lac operon was also a result of a close association between a geneticist,
Francois Jacob and a biochemist, Jacque Monod.
In lac operon (here lac refers to lactose), a polycistronic structural gene is regulated by a common
promoter and regulatory genes.
Such arrangement is very common in bacteria and is referred to as operon
The lac operon consists of
• One regulatory gene (the i gene – here the term i does not refer to inducer, rather it is derived from
the word inhibitor) Thei gene codes for the repressor of the lac operon.
• Three structural genes (z, y, and a).
• The z gene codes for beta-galactosidase (β-gal), which is primarily responsible for the hydrolysis of the
disaccharide, lactose into its monomeric units, galactose and glucose.
• The y gene codes for permease, which increases permeability of the cell to βgalactosides.
• The a gene encodes a transacetylase.
• Hence, all the three gene products in lac operon are required for metabolism of lactose. In most other
operons as well, the genes present in the operon are needed together to function in the same or related
metabolic pathway.
REGULATION OF LAC OPERON IN THE ABSENCE AND PRESENCE OF LACTOSE AS AN INDUCER
Lactose is the substrate for the enzyme beta-galactosidase and it regulates switching on and off of the
operon. Hence, it is termed as inducer.
In the absence of a preferred carbon source such as glucose, if lactose is provided in the growth
medium of the bacteria, the lactose is transported into the cells through the action of permease. The
lactose then induces the operon in the following manner.
a) IN THE ABSENCE OF AN
INDUCER: The repressor of the
operon is synthesised (all-the-time
– constitutively) from the i gene.
The repressor protein binds to the
operator region of the operon and
prevents RNA polymerase from
transcribing the operon. So no
protein is formed by structural genes z,y and a.
b)IN THE PRESENCE OF AN INDUCER:
When inducer such as lactose or allo lactose
is present, the repressor is inactivated by
interaction with the inducer. This allows RNA
polymerase access to the promoter and
transcription proceeds. Structural genes z,y
and a produce β-galactosides,. Permease
and transacetylase respectively.
Essentially, regulation of lac
operon can also be visualised as
regulation of enzyme synthesis by its
substrate.
Regulation of lac operon by repressor is referred to as negative regulation.
HUMAN GENOME PROJECT:
With the establishment of genetic engineering techniques where it was possible to isolate and clone any
piece of DNA and availability of simple and fast techniques for determining DNA sequences, a very
ambitious project of sequencing human genome was launched in the year 1990.
Human Genome Project (HGP) was called a mega project. Human genome is said to have approximately 3
x 109bp, and if the cost of sequencing required is US $ 3 per bp (the estimated cost in the beginning), the
total estimated cost of the project would be approximately 9 billion US dollars.
GOALS OF HGP:
1. Identify all the approximately 20,000-25,000 genes in human DNA;
2. Determine the sequences of the 3 billion chemical base pairs that make up human DNA;
3. Store this information in databases;
4. Improve tools for data analysis;
5. Transfer related technologies to other sectors, such as industries;
6. Address the ethical, legal, and social issues (ELSI) that may arise from the project.
The Human Genome Project was a 13-year project coordinated by the U.S. Department of Energy and the
National Institute of Health. During the early years of the HGP, the Wellcome Trust (U.K.) became a major
partner; additional contributions came from Japan, France, Germany, China and others. The project was
completed in 2003.
Many non-human model organisms, such as bacteria, yeast,
Caenorhabditiselegans (a free living non-pathogenic nematode), Drosophila (the fruit fly), plants
(rice and Arabidopsis), etc., have also been sequenced.
METHODOLOGIES :
The methods involved two major approaches.
i)ESTs:One approach focused on identifying all the genes that are expressed as RNA (referred to as Expressed
Sequence Tags (ESTs).
ii)Sequence annotation: The blind approach of simply sequencing the whole set of genome that contained all
the coding and non-coding sequence, and later assigning different regions in the sequence with functions (a
term referred to as Sequence Annotation).
• For sequencing, the total DNA from a cell is isolated and converted into random fragments of
relatively smaller sizes (recall DNA is a very long polymer and there are technical limitations in
sequencing very long pieces of DNA) and cloned in suitable host using specialised vectors
• The cloning resulted into amplification of each piece of DNA fragment so that it subsequently could be
sequenced with ease.
• The commonly used hosts were bacteria and yeast, and the vectors were called as BAC (bacterial
artificial chromosomes), and YAC (yeast artificial chromosomes).
• The fragments were sequenced using automated DNA sequencers that worked on the principle of a
method developed by Frederick Sanger.
• (Remember, Sanger is also credited for developing method for determination of amino acid
sequences in proteins).
• These sequences were then arranged based on some overlapping regions present in them.
• This required generation of overlapping fragments for sequencing. Alignment of these sequences was
humanly not possible.
• Therefore, specialised computer based programs were developed.
• These sequences were subsequently annotated and were assigned to each chromosome.
• The sequence of chromosome 1 was completed only in May 2006 (this was the last of the 24 human
chromosomes – 22 autosomes and X and Y – to be sequenced). Another challenging task was
assigning the genetic and physical maps on the genome.
• This was generated using information on polymorphism of restriction endonuclease recognition sites,
and some repetitive DNA sequences known as microsatellites (one of the applications of
polymorphism in repetitive DNA sequences shall be explained in next section of DNA fingerprinting).
SALIENT FEATURES OF HUMAN GENOME
(i) The human genome contains 3164.7 million bp.
(ii) The average gene consists of 3000 bases, but sizes vary greatly, with the largest known human gene
being dystrophin at 2.4 million bases.
(iii) The total number of genes is estimated at 30,000–much lower than previous estimates of 80,000 to
1,40,000 genes. Almost all (99.9 per cent) nucleotide bases are exactly the same in all people.
(iv) The functions are unknown for over 50 per cent of the discovered genes.
(v) Less than 2 per cent of the genome codes for proteins.
(vi) Repeated sequences make up very large portion of the human genome.
(vii) Repetitive sequences are stretches of DNA sequences that are repeated many times, sometimes
hundred to thousand times. They are thought to have no direct coding functions, but they shed light on
chromosome structure, dynamics and evolution.
(viii) Chromosome 1 has most genes (2968), and the Y has the fewest (231).
(ix) Scientists have identified about 1.4 million locations where singlebase DNA differences (SNPs – single
nucleotide polymorphism, pronounced as ‘snips’) occur in humans. This information promises to
revolutionise the processes of finding chromosomal locations for disease-associated sequences and
tracing human history.
The repetitive DNA and their polymorphism form the basis of DNA fingerprinting:
DNA fingerprinting involves identifying differences in some specific regions in DNA sequence called as
repetitive DNA, because in these sequences, a small stretch of DNA is repeated many times.
These repetitive DNA are separated from bulk genomic DNA as different peaks during density gradient
centrifugation. The bulk DNA forms a major peak and the other small peaks are referred to as satellite
DNA.
Depending on base composition (A : T rich or G:C rich), length of segment, and number of repetitive
units, the satellite DNA is classified into many categories, such as micro-satellites, mini-satellites etc.
These sequences normally do not code for any proteins, but they form a large portion of human genome.
These sequence show high degree of polymorphism and form the basis of DNA fingerprinting.
Since DNA from every tissue (such as blood, hair-follicle, skin, bone, saliva, sperm etc.), from an individual
show the same degree of polymorphism, they become very useful identification tool in forensic
applications.
Since the polymorphisms are inheritable from parents to children, DNA fingerprinting is the basis of
paternity testing, in case of disputes.
As polymorphism in DNA sequence is the basis of genetic mapping of human genome as well as of DNA
fingerprinting, it is essential that we understand what DNA polymorphism means in simple terms.
Polymorphism (variation at genetic level) arises due to mutations.
New mutations may arise in an individual either in somatic cells or in the germ cells (cells that generate
gametes in sexually reproducing organisms). If a germ cell mutation does not seriously impair individual’s
ability to have offspring who can transmit the mutation, it can spread to the other members of
population (through sexual reproduction).
Allelic sequence variation has traditionally been described as a DNA polymorphism if more than one
variant (allele) at a locus occurs in human population with a frequency greater than 0.01. In simple terms,
if an inheritable mutation is observed in a population at high frequency, it is referred to as DNA
polymorphism.
The probability of such variation to be observed in noncoding DNA sequence would be higher as
mutations in these sequences may not have any immediate effect/impact in an individual’s reproductive
ability.
These mutations keep on accumulating generation after generation, and form one of the basis of
variability/polymorphism.
There is a variety of different types of polymorphisms ranging from single nucleotide change to very large
scale changes.
The technique of DNA Fingerprinting was initially developed by Alec Jeffreys. He used a satellite DNA as
probe that shows very high degree of polymorphism. It was called as Variable Number of Tandem
Repeats (VNTR).
The technique, as used earlier, involved Southern blot hybridisation using radiolabelled VNTR as a probe. It
included
(i) Isolation of DNA,
(ii) Digestion of DNA by restriction endonucleases,
(iii) Separation of DNA fragments by electrophoresis,
(iv) Transferring (blotting) of separated DNA fragments to synthetic membranes, such as nitrocellulose or nylon,
(v) hybridisation using labelled VNTR probe, and (v) Detection of hybridised DNA fragments by autoradiography.
Applications of DNA Fingerprinting
1) To identify the criminals of murder, robbery, rapes, etc.
2) DNA fingerprinting forms the basis of paternity testing since a child inherits polymorphism from both
its parents.
3)It can be used for studying genetic diversity and evolution
VNTR:
A small DNA sequence is arranged tandemly in many copy numbers. The copy number varies from
chromosome to chromosome in an individual. The numbers of repeat show very high degree of
polymorphism. As a result the size of VNTR varies in size from 0.1 to 20 kb.
• Consequently, after hybridisation with VNTR probe, the autoradiogram gives many bands of differing sizes. These bands
give a characteristic pattern for an individual DNA
• It differs from individual to individual in a population except in the case of monozygotic (identical) twins.