CDNA library
● A cDNA library is a combination of cloned cDNA (complementary DNA)
fragments inserted into a collection of host cells, which constitute some
portion of the transcriptome of the organism and are stored as a "library."
● cDNA is produced from fully transcribed mRNA found in the nucleus and
therefore contains only the expressed genes of an organism.
● Similarly, tissue-specific cDNA libraries can be produced.
● In eukaryotic cells, the mature mRNA is already spliced; hence, the cDNA
produced lacks introns and can be readily expressed in a bacterial cell.
● While information in cDNA libraries is a powerful and useful tool since gene
products are easily identified, the libraries lack information about enhancers,
introns, and other regulatory elements found in a genomic DNA library
Construction of cDNA
● cDNA is created from a mature mRNA from a eukaryotic cell with the use of
reverse transcriptase.
● In eukaryotes, a poly-(A) tail (consisting of a long sequence of adenine
nucleotides) distinguishes mRNA from tRNA and rRNA and can therefore be
used as a primer site for reverse transcription
● This has the problem that not all transcripts, such as those for the histone,
encode a poly-A tail
mRNA extraction
● Firstly, mRNA template needs to be isolated for the creation of cDNA libraries.
● Since mRNA only contains exons, the integrity of the isolated mRNA should
be considered so that the protein encoded can still be produced. Isolated
mRNA should range from 500 bp to 8 kb.
● Several methods exist for purifying RNA such as trizol extraction and column
purification.
● Column purification can be done using oligomeric dT nucleotide coated
resins, and features of mRNA such as having a poly-A tail can be exploited
where only mRNA sequences containing said feature will bind.
The desired mRNA bound to the column is then eluted.
.
● Once mRNA is purified, an oligo-dT primer (a short sequence of
deoxy-thymidine nucleotides) is bound to the poly-A tail of the RNA.
● The primer is required to initiate DNA synthesis by the enzyme reverse
transcriptase.
● This results in the creation of RNA-DNA hybrids where a single strand of
complementary DNA is bound to a strand of mRNA.
● To remove the mRNA, the RNAse H enzyme is used to cleave the backbone
of the mRNA and generate free 3'-OH groups, which is important for the
replacement of mRNA with DNA
● DNA polymerase I is then added, the cleaved RNA acts as a primer the DNA
polymerase I can identify and initiate replacement of RNA nucleotides with
those of DNA.
● This is provided by the sscDNA itself by coiling on itself at the 3' end,
generating a hairpin loop. The polymerase extends the 3'-OH end, and later
the loop at 3' end is opened by the scissoring action of S1 nuclease.
● Restriction endonucleases and DNA ligase are then used to clone the
sequences into bacterial plasmids.
The cloned bacteria are then selected, commonly through the use of antibiotic
selection. Once selected, stocks of the bacteria are created which can later be grown
and sequenced to compile the cDNA library.
Applications
● cDNA libraries are commonly used when reproducing eukaryotic genomes, as
the amount of information is reduced to remove the large numbers of
non-coding regions from the library.
● cDNA libraries are used to express eukaryotic genes in prokaryotes
● . Prokaryotes do not have introns in their DNA and therefore do not possess
any enzymes that can cut them out during the transcription process
● . cDNA does not have introns and therefore can be expressed in prokaryotic
cells. cDNA libraries are most useful in reverse genetics, where the additional
genomic information is of less use.
● Additionally, cDNA libraries are frequently used in functional cloning to identify
genes based on the encoded protein's function.
● When studying eukaryotic DNA, expression libraries are constructed using
complementary DNA (cDNA) to help ensure the insert is truly a gene.
cDNA Library vs. Genomic DNA Library
cDNA library lacks the non-coding and regulatory elements found in genomic DNA.
Genomic DNA libraries provide more detailed information about the organism, but
are more resource-intensive to generate and keep.
Genomic library
Cloning DNA
● Cloning DNA by whatever method gives rise to a population of recombinant DNA
molecules.
● Recombinant DNA molecules are often in plasmid vectors or phage vectors.
● These vectors are maintained either in bacterial cells or as phage particles.
● A collection of independent clones is termed a clone bank or library.
● The term genomic library is often used to describe a set of clones representing the entire
genome of an organism.
● Production of such a library is usually the first step in isolating a DNA sequence from an
organism’s genome.
Importance
● A genomic library is a rich resource for the scientist.
● It represents the entire genome of an organism.
● It should contain:
○ all the genes
○ their control sequences
○
Characteristics of a Good Genomic Library
1. Entire Genome Representation
● A genomic library should represent the entire genome of an organism.
2. Overlapping Cloned Fragments
● The genome should be represented as a set of overlapping cloned fragments.
● These overlapping fragments enable the isolation of any sequence in the genome.
3. Sequence-Independent Fragment Generation
● The fragments for cloning should ideally be generated by a sequence-independent
procedure.
● This ensures that fragments are generated at random.
● There is no bias towards any particular sequence.
4. Stable Maintenance
● The cloned fragments should be maintained in a stable form.
● There should be no misrepresentation of sequences caused by:
○ recombination
○ differential replication of the cloned DNAs during propagation of the
recombinants.
5. Practical Reality
● These criteria may seem rather demanding.
● However, the systems available for producing genomic libraries enable these
requirements to be met more or less completely.
First Consideration in Constructing a Genomic Library
● The first consideration is the number of clones required.
Factors Affecting Number of Clones
1. Size of the genome
2. Type of vector used
3. Size of DNA fragments that can be cloned
Example
● Small genome (E. coli) → fewer clones required.
● Complex genome (human genome) → many more clones required.
Calculation of Library Size
In practice, library size can be calculated on the basis of probability of a particular sequence
being represented in the library.
A formula takes into account all the factors and produces a “number of clones” value.
Formula
Where
● N = number of clones required
● P = desired probability of a particular sequence being represented
○ typically 0.95 or 0.99
● a = average size of the DNA fragments to be cloned
● b = size of the genome (expressed in the same units as a)
By using this formula it is possible to:
○ determine the magnitude of the task ahead
○ plan a cloning strategy accordingly
● Some genome sizes and associated library sizes are shown in Table 6.2.
Organism Genome 20 kb 45 kb
size (kb) inser inser
ts ts
Escherichia coli (bacterium) 4.0 × 10³ 6.0 × 2.7 ×
10² 10²
Saccharomyces cerevisiae (yeast) 1.4 × 10⁴ 2.1 × 9.3 ×
10³ 10²
Arabidopsis thaliana (simple higher 7.0 × 10⁴ 1.1 × 4.7 ×
plant) 10⁴ 10³
Drosophila melanogaster (fruit fly) 1.7 × 10⁵ 2.5 × 1.1 ×
10⁴ 10⁴
Strongylocentrotus purpuratus (sea 8.6 × 10⁵ 1.3 × 5.7 ×
urchin) 10⁵ 10⁴
Homo sapiens (human) 3.0 × 10⁶ 4.5 × 2.0 ×
10⁵ 10⁵
Triticum aestivum (hexaploid wheat) 1.7 × 10⁷ 2.5 × 1.1 ×
10⁶ 10⁶
8. Important Notes about the Table
● Number of clones (N) is calculated for probability P = 95% that a given sequence is
represented in the genomic library.
● Genome sizes shown are approximate haploid genome sizes where appropriate.
● Two values of N are shown:
○ 20 kb inserts (replacement vector size)
○ 45 kb inserts (cosmid vectors).
9. These Values are Minimum
Estimates
The calculation assumes:
1. Genome size is known accurately.
2. DNA is fragmented in a totally random manner for cloning.
3. Each recombinant DNA molecule gives rise to a single clone.
4. Efficiency of cloning is the same for all fragments.
5. Diploid organisms are homozygous for all loci.
These assumptions are usually not all valid for a given experiment.
10. Practical Library Size Example
● For a human genomic library, approximately 10⁶ clones or more may be required.
● This ensures reasonable probability of isolating a particular single-copy gene
sequence.
11. Vectors Used for Genomic
Libraries
When dealing with very large libraries, special vectors are required.
1. Phage Vectors
● Often essential for large libraries.
● High cloning capacity and efficiency.
2. Cosmid Vectors
● Can clone DNA fragments up to about 47 kb.
● Seem to be better choice due to larger insert size.
3. Replacement Vectors
● Often used for library construction because:
○ easier to use
○ techniques for screening phage libraries are routine and well characterized.
Advantage
● Important especially for workers new to gene manipulation technology.
Disadvantage
● Only half the cloning capacity compared with cosmids.
Very large DNA fragments can also be cloned using:
● BACs – Bacterial Artificial Chromosomes
● YACs – Yeast Artificial Chromosomes
These vectors allow cloning of very large genomic DNA fragments