Functional Genomics
(MBG419)
Week 3 (5-6.03.2026)
‘The Human Genome Project’
Dr Ömer Faruk Bay
HGP Timeline
Molecular Biology Revolution Set the Stage for the
Human Genome Project (HGP)
DNA DNA Polymerase Chain
Cloning Sequencing Reaction (PCR)
1970s 1977 1983
Discussions on the HGP
1987 1989
1988
Initially Envisioned Plan for HGP
1. Expect to be a 15-year initiative
2. Gain experience with model (i.e., well-studied, experimental) organisms with smaller
genomes before giving full attention to human genome
3. In each case, map (i.e., organize) DNA first and then sequence (i.e., read) DNA
4. Wait to sequence human genome until a new ‘revolutionary’ DNA sequencing method(s)
becomes available – replacing Sanger DNA sequencing
5. Make generating the first sequence of the human genome the signature accomplishment
of the HGP
#1 → An overestimate
#2 → Maintained
#3 → Mostly maintained
#4 → Abandoned
#5 → Absolutely true!
Books to Conceptualize Genomes
Human Genome
Mouse Genome
~3,000,000,000 bp
Fruit Fly Genome
~160,000,000 bp
Nematode Genome
~100,000,000 bp
Yeast Genome
~15,000,000 bp
bp = base pairs
Realities of 1990 (at HGP Launch)
• Scientific community had mixed opinions about HGP
• No detailed start-to-finish plan for executing HGP (i.e., overt expectation
to ‘figure it out along the way’)
• Genomics was a ‘toddler’ field, growing up as a melting pot of scientific
immigrants from other disciplines
• Painfully early days of a functional internet
Implementation of HGP
• International in terms of funding and participants
U.S. Funders:
National Institutes of Health (NIH)
Department of Energy (DOE)
Other Countries:
Some government funders
Some private funders
• Distributed consortium-based (‘team science’) effort
• For studying the human genome: ‘divide & conquer’ strategy
HGP Guided by Periodically Updated Plans
1991-1995 1993-1998 1998-2003
Genomes Organized by Chromosomes
Yeast Nematode Fruit Fly Human
Scale of Genomes, Chromosomes, & Clones
Human Genome
(~3,000 Mb)
Roughly size of
Human Chromosome entire fruit fly or
nematode genome
(~130 Mb)
GATCGTCTAGAATCTC
GATCGTCTAGAATCTC
GATCGTCTAGAATCTC
GATCGTCTAGAATCTC
GATCGTCTAGAATCTC
GATCGTCTAGAATCTC
GATCGTCTAGAATCTC GATCGTCTAGAATCTC
GAGATCTCTGAGAGTC
GAGATCTCTGAGAGTC
GAGATCTCTGAGAGTC
GAGATCTCTGAGAGTC
GAGATCTCTGAGAGTC
GAGATCTCTGAGAGTC
GAGATCTCTGAGAGTC GAGATCTCTGAGAGTC
Larger Clones Smaller Clones
GTGGGAAACTGTGTGA
GTGGGAAACTGTGTGA
GTGGGAAACTGTGTGA
GTGGGAAACTGTGTGA
GTGGGAAACTGTGTGA
GTGGGAAACTGTGTGA
GTGGGAAACTGTGTGA GTGGGAAACTGTGTGA
TGTGACTAGCCACAGT
TGTGACTAGCCACAGT
TGTGACTAGCCACAGT
TGTGACTAGCCACAGT
TGTGACTAGCCACAGT
TGTGACTAGCCACAGT
TGTGACTAGCCACAGT TGTGACTAGCCACAGT
TG TG TG TG TG TG TGTGACTAGCCACAGT TAGGTATTGGGGCATT
TACGTGTGAGAGATGT
TACGTGTGAGAGATGT
TACGTGTGAGAGATGT
TACGTGTGAGAGATGT
TACGTGTGAGAGATGT
TACGTGTGAGAGATGT
TACGTGTGAGAGATGT TACGTGTGAGAGATGT
ATGATGCACCTGACCC
ATGATGCACCTGACCC
ATGATGCACCTGACCC
ATGATGCACCTGACCC
ATGATGCACCTGACCC
ATGATGCACCTGACCC
ATGATGCACCTGACCC ATGATGCACCTGACCC
(~0.5-1.0 Mb) (~0.1-0.2 Mb)
GGGTTTCACTCTCAAC
GGGTTTCACTCTCAAC
GGGTTTCACTCTCAAC
GGGTTTCACTCTCAAC
GGGTTTCACTCTCAAC
GGGTTTCACTCTCAAC
GGGTTTCACTCTCAAC GGGTTTCACTCTCAAC
GACTCACTCCACCTCA
GACTCACTCCACCTCA
GACTCACTCCACCTCA
GACTCACTCCACCTCA
GACTCACTCCACCTCA
GACTCACTCCACCTCA
GACTCACTCCACCTCA GACTCACTCCACCTCA
CC CC CC CC CC CC CCGGTTAGACATACAT CCGGTTAGACATACAT
GAGGCCCACCGCCGCT
GAGGCCCACCGCCGCT
GAGGCCCACCGCCGCT
GAGGCCCACCGCCGCT
GAGGCCCACCGCCGCT
GAGGCCCACCGCCGCT
GAGGCCCACCGCCGCT GAGGCCCACCGCCGCT
GTGCACGTCCACCACC
GTGCACGTCCACCACC
GTGCACGTCCACCACC
GTGCACGTCCACCACC
GTGCACGTCCACCACC
GTGCACGTCCACCACC
GTGCACGTCCACCACC GTGCACGTCCACCACC
Clone-Based Physical Mapping
Chromosome
Clones
Larger Clones (like ‘Chapters’) Smaller Clones (like ‘Pages’)
Clone Contigs
Sequence-Ready Clone Contig Map
Clones highlighted by red rectangles selected for DNA sequencing
Caveat: Note that the book metaphor is imperfect – adjacent clones (i.e., pages)
actually overlap slightly rather than having precise ‘page breaks’ between clones
(i.e., pages).
Shotgun Sequencing
• Shotgun sequencing is a laboratory technique for determining the DNA sequence of an organism’s
genome (or part of the genome). The method involves randomly breaking up the DNA into small
fragments that are then sequenced individually. A computer program looks for overlaps in the DNA
sequences, using them to reassemble the fragments in their correct order to determine the
sequence of the starting DNA.
From NHGRI’s ‘Talking Glossary’
[Link]/genetics-glossary
Subclone Construction
GATCGTCTAGAATCT C
GAGATCTCTGAGAGT C
GTGGGAAACTGTGTG A
TGTGACTAGCCACAG T
Clone DNA
GTGGGAAACTGTGTG A
TACGTGTGAGAGATG T
ATGATGCACCTGACC C
GGGTTTCACTCTCAA C
GACTCACTCCACCTC A
GTGGGAAACTGTGTG A
Prepare Multiple Copies
GAGGCCCACCGCCGC T
GTGCACGTCCACCAC C
GATCGTCTAGAATCT
GATCGTCTAGAATCT
GATCGTCTAGAATCT
GATCGTCTAGAATCT
GATCGTCTAGAATCT
C C C C C
GAGATCTCTGAGAGT
GAGATCTCTGAGAGT
GAGATCTCTGAGAGT
GAGATCTCTGAGAGT
GAGATCTCTGAGAGT
C C C C C
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
A A A A A
TGTGACTAGCCACAG
TGTGACTAGCCACAG
TGTGACTAGCCACAG
TGTGACTAGCCACAG
TGTGACTAGCCACAG
T T T T T
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
A A A A A
TACGTGTGAGAGATG
TACGTGTGAGAGATG
TACGTGTGAGAGATG
TACGTGTGAGAGATG
TACGTGTGAGAGATG
T T T T T
ATGATGCACCTGACC
ATGATGCACCTGACC
ATGATGCACCTGACC
ATGATGCACCTGACC
ATGATGCACCTGACC
C C C C C
GGGTTTCACTCTCAA
GGGTTTCACTCTCAA
GGGTTTCACTCTCAA
GGGTTTCACTCTCAA
GGGTTTCACTCTCAA
C C C C C
GACTCACTCCACCTC
GACTCACTCCACCTC
GACTCACTCCACCTC
GACTCACTCCACCTC
GACTCACTCCACCTC
A A A A A
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
GTGGGAAACTGTGTG
A A A A A
GAGGCCCACCGCCGC
GAGGCCCACCGCCGC
GAGGCCCACCGCCGC
GAGGCCCACCGCCGC
GAGGCCCACCGCCGC
T T T T T
GTGCACGTCCACCAC
GTGCACGTCCACCAC
GTGCACGTCCACCAC
GTGCACGTCCACCAC
GTGCACGTCCACCAC
C C C C C
Randomly Fragment
Subclone Fragments
GATCGTCTAGAATCT C
Shotgun Sequencing Strategy
GAGATCTCTGAGAGT C
GTGGGAAACTGTGTG A
TGTGACTAGCCACAG T
GTGGGAAACTGTGTG A
TACGTGTGAGAGATG T
ATGATGCACCTGACC C
GGGTTTCACTCTCAA C
Clone DNA
GACTCACTCCACCTC A
GTGGGAAACTGTGTG A
GAGGCCCACCGCCGC T
GTGCACGTCCACCAC C
Subclones
Generate Shotgun Sequence Reads
Assemble Sequence Reads into Sequence Contigs
Deduce Sequence
‘Working Draft’ Sequence
Sequence Finishing
Final Sequence
Sequence Reads Like Raindrops on a Sidewalk
Because raindrops fall in random locations, it takes many extra drops in certain areas to ensure that
every portion of the sidewalk gets wet.
The additional consideration for DNA sequencing is that the final accuracy of the sequence depends on
reading every DNA base multiple times (e.g., 30-50 times; called ‘coverage’).
Sequence Assembly Challenges: Repeated Sequences
These fragments are like
DNA sequence reads
As the fragments are aligned to
reconstruct the text, notice that there are
ambiguities
• Imagine you did not know this text.
• If you were to sequence this text the way
we sequence DNA, you would copy the
text, fragment it, and sequence the
• Branches and loops represent alternative
fragments many times over (see right).
assemblies in a complex and often repetitive
genome.
• An actual genome sequence is way more
repetitive and complex than a Dickens novel.
• Requires sophisticated computational tools to
assemble sequence correctly.
Dividing Up Human Genome During HGP
For example…
Challenges of Sequencing the Human Genome
• Human Genome: ~3,000,000,000 nucleotides (bases or base pairs)
• Sanger DNA sequencing Circa 1990: ~500-800 bases per read
• ‘Coverage’ (i.e., number of time each base is read) needed to be high
(e.g., >30-fold) to attain high accuracy
• Roughly half of human genome consists repetitive DNA, much of it
reflecting remnants of transposable elements
Generating the First Human Genome Sequence
Initial HGP Plan Automation & Scale Computational Power
• 6 Countries, 20 Centers, 1000’s of researchers
• ~1,000 bases/second, 24 hours/day, & 7 days/week for ~6 years
• Brute force using Sanger DNA sequencing and massive computational help
Whose Genome Was Sequenced by HGP?
• ~70% of HGP’s human genome
sequence from 1 blood donor of
blended ancestry from Buffalo, NY
• ~30% of HGP’s human genome
sequence from 19 other people
• HGP human genome sequence was
a ‘mosaic’ representation of
multiple people (a ‘reference’)
Bermuda Principles for Data Sharing
• Significant attention to release and sharing of
HGP genome sequence data
• Two seminal meetings in Bermuda in 1996 and
1997
• Landmark agreement for rapid data release and
public access to HGP genome sequence data
• Became known as ‘Bermuda Principles’
• Among the most important legacy of HGP
Two Major Protagonists of HGP Drama
Francis Collins (UVA Alumnus) Craig Venter (UCSD Alumnus)
Two Major Protagonists of HGP Drama
Francis Collins (UVA Alumnus) Craig Venter (UCSD Alumnus)
• Physician (medical geneticist) and scientist • At NIH at beginning of HGP, pioneered use of
expressed-sequence tags (ESTs) as shortcut for
• HGP participant at U. of Michigan before studying genes (sequencing RNA instead of DNA)
becoming Director of NIH’s ‘genome institute’
(succeeding Jim Watson)
• Began patenting human genes at furious pace, arousing
• Became de facto leader of international controversy
consortium of HGP centers sequencing human
genome • Left NIH, founded private research institute, and
became HGP participant
• Later appointed NIH Director by President
Obama • Grew impatient about pace of HGP; left HGP and joined
forces with company that commercialized new
automated instrument for very high-throughput Sanger
DNA sequencing to create Celera Genomics
• Celera Genomics aimed to compete with the HGP in
generating the first human genome sequence and sell
subscriptions for accessing their genomic data
Purported ‘Race’ to Sequence Human Genome
VS
Initial HGP Plan Venter/Celera Plan
‘Clone-by-Clone ‘Whole-Genome
Shotgun Sequencing’ Shotgun Sequencing’
Editorial Aside: Not really a fair ‘race’ since Celera had access to HGP data (but not vice versa)!!!
June 2000: Draft Sequence of Human Genome
Press Coverage of the ‘Race’
Vanity Fair (December 2000)
February 2001: Papers Reporting
Draft Sequence of Human Genome
HGP Paper Venter/Celera Paper
After June 2000 Announcement & February
2001 Publications
• Venter/Celera could not fully assemble the human genome
sequence and relied on the publicly available data to resolve many
of the difficult regions; had little interest in improving (i.e., ‘finishing’)
sequence beyond ‘working draft’ quality
• HGP focused on improving the human genome sequence from a
‘working draft’ to high-quality ‘finished’
• Celera’s business plan to sell subscription access to the human
genome sequence eventually failed
• Venter moved on to various other endeavors
Generating the First Human Genome Sequence
+ =
Initial HGP Plan Venter/Celera Plan Ultimate HGP Plan
April 25, 2003: HGP Completion
• National DNA Day established
• HGP completion & 50th anniversary of
discovery of DNA’s double-helical structure
Highlight Features of HGP
• Completed ahead of schedule (13 years) and underbudget
• Signature accomplishment was generation of an extremely high-quality
sequence for >90% (‘near-complete’ or ‘essentially complete’) of human
genome
• Cost of generating first human genome sequence by HGP: ~$1 billion
• The ‘race’ between HGP and Venter/Celera melted away after
announcement of draft human genome sequence in 2000
• Similarly, the initial concerns about the HGP from some parts of the
scientific community largely melted away
• HGP set the field of genomics into a trajectory of widespread
dissemination across biology, medicine, and society
Epilogue: A Truly Complete Human
Genome Sequence
• HGP produced a high-quality human genome sequence, but it only
accounted for 92% of the human genome
• Remaining 8% was not ‘readable’ using the then-available methods for
DNA sequencing, but those regions are important for structural
(centromere and telomeres) and medical reasons
• Several new ‘revolutionary’ methods for DNA sequencing have been
developed over the last ~20 years
• These new methods plus better computational approaches set the stage
for a new group of researchers to (finally) generate a truly complete
sequence of the human genome in 2022
2022: A Truly Complete Human Genome Sequence
Take-Home Messages
• HGP: 1990-2003
• HGP used a map-first, sequence-second strategy to study the human genome
• HGP used Sanger DNA sequencing – not a revolutionary new DNA sequencing method
• Sequencing the human genome was particularly difficult because of its large size, complexity, and
extensive amounts of repetitive regions
• Genome sequence assembly was (and remains) a major challenge; repetitive regions present a
particular obstacle to accurately assembling genome sequences
• Venter/Celera pursued a whole-genome sequencing strategy and tried to build a business selling
access to their data; both efforts fell short of expectations
• Ultimately, the HGP completed the task of generating the first high-quality ‘essentially complete’
sequence of the human genome; 19 years later (in 2022), a truly complete (‘telomere-to-telomere’)
human genome sequence was finally generated
Scale of the Human Genome Sequence
[Link]
To Learn More about HGP…
[Link]/HGP