0% found this document useful (0 votes)
19 views7 pages

Paper1 Fu

The document reports on recent segmental duplications in the human genome and their role in gene evolution. It finds that 6.1% of human genes contain exons that originated from segmental duplications within the last 40 million years. Certain gene families related to immunity, membrane interactions, and development are particularly enriched among recently duplicated genes. This suggests segmental duplication has contributed to adaptations and protein diversity in primate evolution.

Uploaded by

api-3700537
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views7 pages

Paper1 Fu

The document reports on recent segmental duplications in the human genome and their role in gene evolution. It finds that 6.1% of human genes contain exons that originated from segmental duplications within the last 40 million years. Certain gene families related to immunity, membrane interactions, and development are particularly enriched among recently duplicated genes. This suggests segmental duplication has contributed to adaptations and protein diversity in primate evolution.

Uploaded by

api-3700537
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

REPORTS

compared with chromosomal duplication con- sponds to duplications that have emerged over 11. T. H. Shaikh et al., Hum. Mol. Genet. 9, 489 (2000).
tent (R2 ⫽ 0.16). The correlation was due to the last ⬃40 million years of human evolution 12. J. A. Bailey et al., Am. J. Hum. Genet. 70, 83 (2002).
13. L. Edelmann, R. K. Pandita, B. E. Morrow, Am. J. Hum.
intrachromosomal duplications (fig. S5; R2 ⫽ (24). Gene duplication followed by functional Genet. 64, 1076 (1999).
0.20; P ⫽ 0.04; F test) and was absent for specialization has long been considered a major 14. R. Mazzarella, D. Schlessinger, Genome Res. 8, 1007
interchromosomal duplications (R2 ⫽ 0.002). evolutionary force for gene innovation (25). (1998).
The three most gene-rich chromosomes showed Therefore, these genes embedded within recent 15. J. R. Lupski, Trends Genet. 14, 417 (1998).
16. S. T. Sherry et al., Nucleic Acids Res. 29, 308 (2001).
high levels of duplication, and the seven most genomic duplications may be considered excel- 17. K. Chen et al., Nature Genet. 17, 154 (1997).
gene-poor chromosomes were among the least lent candidates for adaptations specific to pri- 18. S. L. Christian, J. A. Fantes, S. K. Mewborn, B. Huang,
duplicated chromosomes. mate evolution. D. H. Ledbetter, Hum. Mol. Genet. 8, 1025 (1999).
To determine what role recent segmental 19. D. E. Jenne et al., Am. J. Hum. Genet. 69, 516 (2001).
20. T. Kuroda-Kawaguchi et al., Nature Genet. 29, 279
duplications have played in current gene evo- References and Notes (2001).
lution, we characterized the gene content in 1. International Human Genome Sequencing Consor- 21. H. C. Mefford, B. J. Trask, Nature Rev. Genet. 3, 91
our filtered set of duplicated genomic se- tium, Nature 409, 860 (2001). (2002).
2. J. A. Bailey, A. M. Yavor, H. F. Massa, B. J. Trask, E. E. 22. J. Guy et al., Hum. Mol. Genet. 9, 2029 (2000).
quence. We analyzed a highly curated set of Eichler, Genome Res. 11, 1005 (2001). 23. M. Ashburner et al., Nature Genet. 25, 25 (2000).
13,351 mRNAs assigned to the human ge- 3. J. C. Venter et al., Science 291, 1304 (2001). 24. W. Li, Molecular Evolution (Sinauer Associates, Sun-
nome assembly (RefSeq, [Link]- 4. S. Ohno, U. Wolf, N. Atkin, Hereditas 59, 169 (1968). derland, MA, 1997).
.gov/LocusLink/[Link]). We partitioned 5. E. E. Eichler, Trends Genet. 17, 661 (2001). 25. S. Ohno, Evolution by Gene Duplication (Springer-
6. P. Stankiewicz, J. R. Lupski, Trends Genet. 18, 74 Verlag, Berlin, 1970).
exons from each gene into a unique or dupli- (2002). 26. We thank L. Christ, M. Eichler, and U. Neuss for
cated sequence on the basis of their map 7. E. E. Eichler, Genome Res. 11, 653 (2001). technical assistance, and H. Willard, J. Nadeau, T.
position (⬎90% sequence identity). We iden- 8. See supporting data on Science Online. Hassold, D. Locke, and J. Horvath for helpful com-
9. V. E. Cheung et al., Nature 409, 953 (2001). ments. Supported by NIH grants GM58815 and
tified a total of 7777 exons as being tran- 10. The sequence and underlying test statistics for all HG002318 and U.S. Department of Energy grant
scribed from recently duplicated sequence, duplicated regions of the genome are available at ER62862 (E.E.E.), NIH Career Development Program
corresponding to 6.1% of all RefSeq exons [Link] WGAC com- in Genomic Epidemiology of Cancer (CA094816)
parisons of the human genome assembly (UCSC, and Medical Scientist Training Grant ( J.A.B.), the
(128,467). This is slightly greater than the August freeze, 2001; [Link] were W. M. Keck Foundation, and the Charles B. Wang
genomic representation of segmental duplica- done as described (2). The WSSD-filtered set of Foundation.
tion (5.2%), which confirms that gene-poor WGAC duplications can be interactively searched
([Link] This includes Supporting Online Material
regions have not been preferentially duplicat- extracted sequence files, the actual alignments, the [Link]/cgi/content/full/297/5583/1003/
ed. In many cases, a complete complement of location of the alignments within the assembly, and DC1
exons was not duplicated. These incomplete whole-chromosomal views comparing WGAC and Materials and Methods
WSSD duplication patterns. An updated WSSD based Tables S1 to S7
duplicated genes were often found adjacent to
on the analysis of 39,298 clones from April 2002, Figs. S1 to S5
other duplicated cassettes that originated detecting an additional 36 Mb of duplicated se-
from elsewhere in the genome. By comparing quence, is also available. 20 March 2002; accepted 7 June 2002
our data with human expressed sequence tag
databases, we found evidence for “chimeric”
or fusion transcripts that emerged from the
physical juxtaposition of incomplete segmen- Predictive Identification of
tal duplications. Although the mechanism for
recent segmental duplications is not under- Exonic Splicing Enhancers in
stood, the existing data suggest the process
may play a role in exon shuffling associated
with expanding protein diversity. A complete
Human Genes
list of all genes with one more exons within William G. Fairbrother,1,2* Ru-Fang Yeh,1* Phillip A. Sharp,1,2
duplicated genomic sequence is available (8). Christopher B. Burge1†
To further assess whether specific kinds of
genes or biological processes have been prefer- Specific short oligonucleotide sequences that enhance pre-mRNA splicing when
entially duplicated, we compared all RefSeq present in exons, termed exonic splicing enhancers (ESEs), play important roles
mRNAs on the basis of their INTERPRO pro- in constitutive and alternative splicing. A computational method, RESCUE-ESE,
tein domain classification (Table 2) (table S7) was developed that predicts which sequences have ESE activity by statistical
(23). In this analysis, we considered a gene analysis of exon-intron and splice site composition. When large data sets of
duplicated only if all its exons were contained human gene sequences were used, this method identified 10 predicted ESE
within a duplicated genomic region. Our anal- motifs. Representatives of all 10 motifs were found to display enhancer activity
ysis suggests a nonrandom distribution of seg- in vivo, whereas point mutants of these sequences exhibited sharply reduced
mental duplications within the proteome. Genes activity. The motifs identified enable prediction of the splicing phenotypes of
associated with immunity and defense (natural exonic mutations in human genes.
killer receptors, defensins, interferons, serine
proteases, cytokines), membrane surface inter- Human genes are generally transcribed as must be precisely removed and flanking ex-
actions (galectins, HLA, lipocalins, carcinoem- much longer precursors, typically tens of ki- ons precisely ligated to create the mRNA that
bryonic antigens), drug detoxification (cyto- lobases in length, from which large introns will direct protein synthesis. Sequences
chrome P450), and growth/development (soma- around the splice junctions—the 5⬘ and 3⬘
totropins, chorionic gonadotropins, pregnancy- splice sites (5⬘ss and 3⬘ss)—are clearly im-
1
Department of Biology, 2Center for Cancer Research,
specific glycoproteins) were particularly Massachusetts Institute of Technology, Cambridge,
portant for splice site recognition. However,
enriched. It should be emphasized that our gene MA 02139, USA. these signals appear to contain only about
analysis is restricted to genomic segments that *These authors contributed equally to this work.
half of the information required for exon and
show ⱖ90% sequence identity. On the basis of †To whom correspondence should be addressed. E- intron recognition in human transcripts (1).
neutral expectation of divergence, this corre- mail: cburge@[Link] The sequence or structure context in the vi-

[Link] SCIENCE VOL 297 9 AUGUST 2002 1007


REPORTS
cinity of the 5⬘ss and 3⬘ss motifs is known to splicing at nearby sites (5), are an important through the analysis of disease alleles (6), by
play an important role in splice site recogni- component of this context. site-directed mutagenesis of minigene con-
tion (2– 4). ESE sequences, which enhance Exonic enhancers have been identified structs, and by protocols based on SELEX
(Systematic Evolution of Ligands by EXpo-
nential enrichment) to identify sequences
with enhancer activity from a pool of random
sequences (7–11). These methods initially
characterized ESEs as purine-rich sequences,
but additional classes of AC-rich motifs and
pyrimidine-rich motifs have since emerged
(7, 10).
Our strategy for identifying human ESE
sequences was to first develop a statistical/
computational method to predict the ESE activ-
ity of oligonucleotide sequence motifs, to apply
this method to large data sets of human genom-
ic sequences, and then to test representatives of
each predicted motif by means of an in vivo
splicing assay. At the heart of this approach is a
sequence analysis method that we call
RESCUE (Relative Enhancer and Silencer
Classification by Unanimous Enrichment).
RESCUE identifies the set of oligonucleotide
motifs that enhance or repress a particular bio-
chemical process; it consists of four steps: (i)
Identify two or more statistical “attributes” that
should be manifested by sequences that en-
hance (or, alternatively, repress) the biochemi-
cal activity of interest. (ii) Use a statistical
power calculation to determine an oligonucleo-
tide “word” size k appropriate for the amount of
data available. Then represent all possible oli-
gonucleotides of size k by points in a multidi-
mensional space, the axes of which represent
the attributes chosen in the previous step. (iii)
Define a region in this space corresponding to
“unanimous enrichment” (i.e., significantly
high values of all of the chosen attributes) and
identify clusters of similar sequences that fall in
this region. (iv) Align the sequences in each
cluster to produce motifs, and test representa-
tive sequence(s) from each motif and appropri-
ate point mutants with the use of a suitable
functional assay.
A large body of work suggests that ESEs
are located in the general vicinity of splice
sites (12). Unlike transcriptional enhancers,
ESEs function in a strongly position-depen-
dent manner, enhancing splicing when
Fig. 1. Schematic of RESCUE-ESE approach. Exon-intron structures of human genes are derived by
spliced alignment of cDNAs to the assembled genomic sequence, and splice sites are scored as
present downstream of a 3⬘ss and/or upstream
described (17). Values of ⌬EI (scaled difference in frequency between exons and introns) and ⌬WS of a 5⬘ss (13), but often repressing splicing
(scaled difference in frequency between weak and strong exons) are calculated as described for when present in intronic locations (14, 15).
each of the 4096 possible hexanucleotides (17). Each hexamer is then represented by a colored These observations suggest that, as one at-
letter at the point (⌬EI, ⌬WS) in the scatterplot. The letters are chosen to reflect the base tribute, ESE sequences should be strongly
composition of the hexamer according to IUPAC nomenclature (e.g., hexamers containing only A selected for in constitutively spliced exons
and G are represented by the letter “r”). Hexamers containing homonucleotide runs of three or
more bases (e.g., AAA) are represented by capital letters, all other hexamers by lowercase letters.
and generally avoided in intronic sequences
Each letter is colored proportional to the relative content of A (red), C (green), G (blue), and T near splice sites.
(black) of the hexamer. Hexamers (6mers) satisfying ⌬EI ⬎ 2.5 and ⌬WS ⬎ 2.5 (upper right portion Moreover, ESEs can compensate for the
of first quadrant) are predicted to have ESE activity. As a test of ESE activity, a 19-base “extended presence of “weak” (nonconsensus) 5⬘ or 3⬘
exemplar” sequence containing the hexamer in its natural context in a weak exon is chosen and splice signals in exons, and strengthening of
inserted into the SXN splicing reporter construct as indicated. SXN is a ␤-globin– derived minigene the splice sites of an enhancer-dependent
with deleted translation start codon. A point mutant predicted to disrupt ESE activity is also chosen,
generally the single-base mutant that is farthest to the left and below the predicted ESE hexamer
exon generally eliminates enhancer depen-
in the scatterplot. Transient transfection of the reporter construct followed by quantitative RT-PCR dence (16). Therefore, we conjecture that ex-
with flanking primers is used to assay inclusion of the test exon for the candidate ESE and its ons with nonconsensus splice sites (“weak
mutant. exons”) are under much stronger selective

1008 9 AUGUST 2002 VOL 297 SCIENCE [Link]


REPORTS
pressure to retain ESEs than are exons with as candidate 3⬘ESEs. These two sets overlap represent just three distinct classes, each com-
consensus splice sites (“strong exons”), re- fairly extensively, with 63 of the 103 predict- prising the union of the pair of similar hexamer
sulting in a significantly higher frequency of ed 5⬘ESEs also contained in the set of pre- clusters (17). The total number of distinct can-
ESEs in weak exons than in strong exons. dicted 3⬘ESEs, which suggests that many en- didate enhancer motifs identified by RESCUE-
Available full-length cDNA sequences hancers may be capable of acting at both ESE was therefore 10.
were aligned to the assembled human ge- splice sites [e.g., (13)]. The total number of In the final step of the RESCUE proce-
nome by means of the spliced alignment al- hexamers predicted to display either 5⬘ or 3⬘ dure, representatives of these candidate en-
gorithm that is part of the “Genoa” gene ESE activity was 238 out of the 4096 possible hancer motifs were tested for ESE activity in
annotation script (17). Reliable full-length hexamers, about 6% of the total, consistent a splicing reporter construct. For each cluster
alignments were obtained with this approach with the notion that ESEs are quite common. of predicted ESEs, a representative hexamer
for 4817 human genes containing 31,463 in- In step three of the RESCUE procedure, was chosen—referred to as the “exemplar” of
trons and 28,933 internal exons. Position- predicted 5⬘ESE and 3⬘ESE hexamers were the class. To place each exemplar hexamer in
specific log-odds score matrices were then clustered on the basis of sequence similarity, its natural context, we screened our human
used to score the 5⬘ss and 3⬘ss of these exons, and the hexamers in each cluster were multiply spliced gene database for an occurrence of
and the distributions of 5⬘ss and 3⬘ss scores aligned using CLUSTALW (18) to identify each exemplar hexamer in a weak 5⬘ exon
were used to partition exons into categories candidate enhancer motifs (fig. S3) (17). This (bottom 10% of 5⬘ss scores) or weak 3⬘ exon
on the basis of the strength of their splice procedure yielded a total of five 5⬘ESE motifs (bottom 10% of 3⬘ss scores), as appropriate.
sites: “weak 5⬘ exons” (bottom 25% of 5⬘ss (Fig. 2A) and eight 3⬘ESE motifs (Fig. 2B). A slightly longer region of sequence centered
scores), “strong 5⬘ exons” (top 25% of 5⬘ss Three of the five 5⬘ESE motifs—5A, 5B, and on the exemplar—referred to as the “extend-
scores), with “weak 3⬘ exons” and “strong 3⬘ 5C—are significantly similar to 3⬘ESE motifs ed exemplar”—was then chosen from this
exons” defined analogously. 3G, 3A, and 3D, respectively, so the three pairs exon and inserted into the reporter construct
Application of the RESCUE-ESE method 5A/3G, 5B/3A, and 5C/3D were considered to described below. The extended exemplar
to this set of human genes is illustrated in Fig.
1. A power calculation dictated the use of a
word size of six nucleotides, which is com-
parable in size to the binding sites of many
known RNA binding factors (17). In step
two, each of the 4096 oligonucleotides of
length six was assigned two scores: ⌬EI, the
scaled difference between the frequency of
occurrence of the hexamer in exons and the
frequency of occurrence near splice sites in
introns (scaled in standard deviation units);
and ⌬5WS, the scaled difference between the
frequency of occurrence of the hexamer in
weak 5⬘ exons and its frequency in strong 5⬘
exons (SD units), with ⌬3WS defined analo-
gously for weak 3⬘ exons versus strong 3⬘
exons. Each hexamer was then represented
by a point in the plane with coordinates (⌬EI,
⌬5WS) for identification of sequences that
enhance 5⬘ss recognition (5⬘ESEs) (Fig. 2A).
Alternatively, each hexamer was represented
by the point (⌬EI, ⌬3WS) for identification
of sequences that enhance 3⬘ss recognition
(3⬘ESEs) (Fig. 2B). A statistical significance
threshold of 2.5 standard deviations above
the mean (corresponding to a P value of
⬃0.01) was then applied to each axis inde-
pendently; that is, any hexamer for which
both ⌬EI ⬎ 2.5 and ⌬5WS ⬎ 2.5 is predicted
to be a 5⬘ESE, and any hexamer with both
⌬EI ⬎ 2.5 and ⌬3WS ⬎ 2.5 is predicted to be
a 3⬘ESE (hexamers in the upper right portion
of the first quadrant in the scatterplots). The
requirement that each hexamer exceed
thresholds in two separate dimensions, both Fig. 2. RESCUE-ESE prediction of 5⬘ and 3⬘ ESEs in human genes. (A) Scatterplot for prediction of
with P ⬃ 0.01, represents essentially a Bon- 5⬘ESE activity. Hexamers are represented by colored letters as described in Fig. 1. Simplified
ferroni-type correction for multiple compari- dendrogram shows clustering of 5⬘ESE hexamers (total of 103 hexamers with ⌬EI ⬎ 2.5 and
sons: Because 4096 different tests are being ⌬5WS ⬎ 2.5) into five clusters of four or more hexamers. (B) Scatterplot for prediction of 3⬘ESE
performed, the combined P value is set to activity. Simplified dendrogram shows clustering of 3⬘ESE hexamers (total of 198 hexamers with
⌬EI ⬎ 2.5 and ⌬3WS ⬎ 2.5) into eight clusters of four or more hexamers. Complete dendrograms
⬃(0.01)2 ⫽ 10⫺4, giving an expectation of of all hexamers are shown in fig. S3. The aligned sequences in each cluster are represented as
less than one false positive hexamer. Pictograms ([Link] Cluster labels (e.g., 3B, 5A/3G) are listed to the
These criteria identified 103 different hex- right of each Pictogram, with the total number of hexamers in the cluster indicated in parentheses.
amers as candidate 5⬘ESEs and 198 hexamers Clustering and alignment were performed as described (17).

[Link] SCIENCE VOL 297 9 AUGUST 2002 1009


REPORTS
sequences comprise the 19-base region ex- tained additional purine-rich hexamers reduced levels of exon inclusion in the
tending from six bases 5⬘ of the exemplar overlapping the central GAAGAA hex- context of GAAGAA.3. On the other hand,
hexamer to seven bases 3⬘ of the exemplar amer, which were also predicted to have the mutation A5⬎C (Fig. 3, M3) is predict-
hexamer (Fig. 1). enhancer activity by RESCUE-ESE (indi- ed to preserve ESE activity because it con-
The splicing enhancer activity of each ex- cated by the vertical blue bars in Fig. 3). verts GAAGAA to GAAGCA, another pre-
tended exemplar sequence was then assessed by Next, the mutation G4⬎T was introduced dicted ESE hexamer, and this mutation
measuring its ability to “rescue” splicing of into each extended exemplar [i.e., each cen- slightly increases exon inclusion in the con-
exon 2 of the reporter construct, pSXN (7). tral GAAGAA hexamer was mutated to text of GAAGAA.3 (Fig. 3). These data
SXN exon 2 is only 32 bases long, including the GAATAA, a hexamer that falls far “south- anecdotally suggest that RESCUE-ESE can
19-base insert. Previously, this exon was ob- west” of GAAGAA in the scatterplots and is accurately predict which mutations will
served to be predominantly skipped for most predicted to lack ESE activity (see Fig. 4, disrupt the enhancing activity of an ESE;
random insert sequences tested. This failure to motif 5C/3D)]. This mutation also disrupts some evidence for this conclusion is dis-
be included is reversed when the exon is length- many or all (for GAAGAA.3) of the overlap- cussed below.
ened, when a splicing enhancer is present, or ping RESCUE-ESE hexamers. As predicted, To assess the degree to which different
when the 5⬘ss, the branch point, or the poly- this mutation produced sharply reduced lev- exemplar hexamers from the same cluster
pyrimidine tract is improved (fig. S1) (19 –21). els of inclusion in each of the three contexts, would have similar ESE activity, we chose a
Because strengthening either the 5⬘ss or 3⬘ss ranging from ⬃5% to ⬃30% of the wild-type quite different exemplar, AGAAAC, from the
consensus sequence of SXN exon 2 or inserting level (Fig. 3). Taken together, these data sug- same 5C/3D cluster as GAAGAA. The ex-
an ESE causes exon inclusion, we reasoned that gest that different occurrences of the same tended exemplar AGAAAC.1 also displayed
this exon would be a suitable reporter system exemplar tend to be qualitatively similar in ESE activity in the range observed for the
for testing the activity of candidate 5⬘ESEs as their ability to enhance splicing and in their different extended exemplars of GAAGAA
well as candidate 3⬘ESEs. response to specific point mutations, but that (Fig. 3). However, the mutation G2⬎T, pre-
It was of particular interest to assess the the precise level of ESE activity depends dicted to disrupt the activity of AGAAAC,
ability of the RESCUE approach to predict on local sequence context. Another muta- gave only a moderate (⬃27%) reduction in
ESE-disrupting mutations. Therefore, for tion predicted to disrupt ESE activity of exon inclusion, from ⬃75% to ⬃55%. This
each exemplar hexamer, a single-base mutant GAAGAA, A2⬎T (Fig. 3, M2), also gave remaining ESE activity might be attributable
was chosen that was predicted to lack en-
hancer activity [i.e., did not fall in the ex-
treme upper right (“unanimous enrichment”)
region of the scatterplot]. Typically, the sin-
gle-point mutant farthest “southwest of ” (to
the left and below) the exemplar in the scat-
terplot was chosen. A “mutant” extended ex-
emplar sequence containing just this single
base change was then generated for each
extended exemplar and inserted into the same
cloning site in the SXN minigene. Constructs
containing the extended exemplars and mu-
tants were transiently transfected into HeLa
cells, and the splicing phenotype was assayed
by quantitative reverse-transcription poly-
merase chain reaction (RT-PCR) (see fig. S2
for protocol and quantitation curves).
An initial set of experiments evaluated
the robustness of the approach with respect
to differences in the local context of the
exemplar hexamer. For this purpose we
focused on a representative hexamer,
GAAGAA, chosen from the large purine-
rich 5C/3D cluster of predicted enhancers Fig. 3. Analysis of ESE activity for predicted enhancers of class 5C/3D. Upper panel: Extended
(Fig. 2). The consensus sequences for these exemplar sequences for three occurrences of the GAAGAA exemplar and one occurrence of the
clusters and the chosen exemplar are simi- AGAAAC exemplar. All extended exemplars derive from arbitrarily selected occurrences of the
lar to the classical “GARGAR” enhancer exemplar in human exons with weak splice sites, as described in the text. Gene name and exon
(R represents either purine nucleotide, A or number are listed above each sequence. GenBank accession numbers for the mRNAs are as follows:
G). Occurrences of GAAGAA were identi- XM_046769 (GAAGAA.1), XM_010365 (GAAGAA.2), AF212232 (GAAGAA.3), and BC020651
(AGAAAC.1). Predicted ESE hexamers in each extended exemplar are indicated by blue bars above
fied in three exons with weak splice sites, the sequence. Point mutations introduced into these sequences are shown in red, with predicted
generating extended exemplar sequences GAA- ESE hexamers in the mutant sequence shown by blue bars below the sequence. Each mutant is
GAA.1, GAAGAA.2, and GAAGAA.3, which labeled by a red M if the mutation is predicted to disrupt ESE activity, or by a blue M if the mutant
lack appreciable similarity other than the sequence is predicted to retain ESE activity. Total RNA extracted from HeLa cells was amplified by
shared hexamer GAAGAA (Fig. 3). All three RT-PCR after transient transfection with the SXN reporter containing the indicated insert. Radio-
extended exemplars conferred high levels of labeled products were analyzed by polyacrylamide gel electrophoresis and visualized using a
phosphorimager. (Representative autoradiographs are shown in fig. S4.) Bottom panel: Percent
inclusion on the test exon, ranging from inclusion for each construct was calculated as the ratio of the intensity of the upper band (including
⬃50% for GAAGAA.3 to ⬃70% for GAA- exon 2) to the sum of the intensities of the upper and lower bands. All transfections were
GAA.1 (see fig. S4 for representative gels). performed at least twice. The height of the colored bar indicates the average of all measurements;
All three of these extended exemplars con- horizontal black lines indicate the minimum and maximum inclusion values observed.

1010 9 AUGUST 2002 VOL 297 SCIENCE [Link]


REPORTS
to the retention of two predicted ESE hexa- sequence gave a significantly higher level of than 2000 alternative (skipped) exons than in
mers in the mutated AGAAAC.1 sequence inclusion than the mutant (blue bar higher our database of constitutively spliced exons
(Fig. 3), although this was not tested. This than red bar), motif 3F being the only excep- (22); this finding suggests that the motifs we
example underscores the difficulty in inter- tion (see fig. S5 for representative gels). have identified are involved in recognition of
preting the results of mutations in sequences These results demonstrate the effectiveness both constitutively and alternatively spliced
containing additional predicted ESE hexa- of RESCUE-ESE for prediction of the effects exons.
mers. To rigorously test the predictions of the of single base changes on ESE activity. The Some sequences that display ESE activity
RESCUE-ESE method, we used sequences different point mutant sequences exhibited were missed by the RESCUE method in its
specifically chosen to contain exactly one varying levels of inclusion. Mutant 3F gave current form. For example, mutant 3F and
predicted ESE hexamer (or zero, in the case comparable inclusion to wild-type 3F en- one of the three predicted “neutral” sequenc-
of mutant sequences) in all other experiments hancer. Three other mutants, 3C, 3E, and 3H, es tested (fig. S5) displayed enhancer activ-
reported here. gave about two-thirds the level of inclusion ity but did not contain any RESCUE-pre-
One exemplar and a corresponding ex- of the wild-type sequence, indicating that dicted ESE hexamers. Analysis of the three
tended exemplar were chosen from each of ESE activity had been only partially im- predicted neutral sequences—19-base seg-
the 10 motifs for testing in the reporter con- paired. On the other hand, the remaining six ments that lack predicted ESE hexamers
struct. Only 19-nucleotide oligomers contain- mutants all had 10 to 50% of the wild-type chosen from exons with weak splice sites—
ing a single RESCUE-ESE–predicted hex- level of inclusion. In absolute terms, these six suggests a possible modification of the cut-
amer in the middle were considered (table mutants had inclusion levels below 20% and offs used in the RESCUE-ESE protocol.
S1). Although this restriction might result in often less than 10%, comparable to that seen Specifically, it was found that neutral se-
a bias toward selection of weaker enhancers, by others for typical random inserts in this quence N3 contained a hexamer that was
it was considered essential in order to avoid context (7). close to the cutoff for ESEs: The hexamer
complications in interpreting the splicing In a large set of human exons, slightly CTACGC had ⌬EI ⫽ 16.9 (⬎⬎ 2.5) and
phenotypes of sequences containing overlap- more than 10% of all the hexanucleotides ⌬5WS ⫽ 2.2, just below the cutoff. By
ping or adjacent predicted ESEs. For each were found to match RESCUE-ESE hexa- contrast, no hexamer in neutral sequence
motif, a single-base mutant predicted to dis- mers (22), often in overlapping clumps as in N1 or N2 had both ⌬EI ⬎ 2.5 and ⌬5WS or
rupt the ESE activity of the exemplar hex- Fig. 3, suggesting that ESEs are very com- ⌬3WS ⬎ 1.5, and no hexamer in any of the
amer was chosen as described above, intro- mon in human genes. Counting each overlap- neutral sequences had ⌬5WS or ⌬3WS ⬎
duced into the extended exemplar sequence, ping clump as a single enhancer, we found an 2.5. These and other data (22) suggest that
and cloned into the reporter construct. The 10 average of 5.2 predicted enhancers per exon, altering the cutoffs used in RESCUE-ESE,
predicted enhancer and mutant constructs with most exons containing between three perhaps by increasing the ⌬EI cutoff while
were transiently transfected and assayed for and seven ESEs (20th and 80th percentiles, simultaneously reducing the ⌬WS cutoff to
splicing as before (Fig. 4). respectively). The hexamers in each cluster 1.5 or 2, might result in improved detection
All 10 of the predicted enhancers (blue typically occurred more frequently in exons of ESEs.
bars) displayed ESE activity in the reporter than in introns by a factor of 1.5 to 2 and A database of published mutationally char-
system, ranging from weakly enhancing more frequently in weak exons than in strong acterized natural ESE sequences was construct-
(⬃20% inclusion for 5B/3A and 3B) to exons by a factor of 1.3 to 1.4. The average ed, and these sequences were searched for
strongly enhancing (⬃60 to 80% inclusion, frequencies of hexamers in each of the 10 occurrences of the hexamers in each cluster
for 5D and 3E). In addition, for 9 of 10 RESCUE-ESE motif clusters were compara- (tables S2 to S4). Five of the RESCUE-ESE
classes of enhancer tested, the predicted ESE ble or slightly lower in a database of more clusters (5C/3D, 5E, 3C, 3E, and 3F) resemble

Fig. 4. Analysis of ESE activity for 10 classes of predicted enhancers. HeLa each construct, calculated as in Fig. 3. (Representative autoradiographs
cells were transfected with SXN splicing reporter construct containing are shown in fig. S5.) All transfections were performed at least twice. The
inserts representing all 10 classes of predicted ESEs and point mutants of height of the colored bar (blue for predicted ESE, red for mutant
these sequences. The extended exemplar sequences used are listed in predicted to disrupt ESE activity) indicates the average of all measure-
table S1. Upper panel: schematic representing ⌬EI and ⌬WS values for ments; horizontal black lines indicate the minimum and maximum in-
each tested exemplar hexamer (blue E) and point mutant hexamer (red clusion values observed. The predicted ESE hexamer and point mutant
M) from Fig. 2A or 2B, as appropriate. The label of the predicted ESE sequence are shown below the blue and red bars, respectively, with the
cluster from Fig. 2 is indicated above. Lower panel: Percent inclusion for mutated base shown in the corresponding color.

[Link] SCIENCE VOL 297 9 AUGUST 2002 1011


REPORTS
Fig. 5. Correlation between predicted
ESEs and exon skipping mutations in
human HPRT gene. (A) Exon skipping
mutations were analyzed in terms of
the set of hexamers affected; mutation
2 (G88⬎T) is shown here as an exam-
ple. All hexamers affected by the mu-
tation are shown, with matches to
RESCUE-predicted ESEs shown in blue
and represented by a plus sign. The
location of the mutation is indicated by
a red arrow. (B) Summary of human
HPRT gene mutations known to cause
exon skipping [from (23, 24)]. Base
changes that occurred within five nu-
cleotides of a splice junction were ex-
cluded, as they may alter the splice site
signals. The first column lists the na-
ture of the mutation (del ⫽ deletion,
X⬎Y ⫽ substitution of base Y for base
X), with coordinates listed relative to
the translation start site of the HPRT
cDNA (GenBank accession number
NM_000194). The last two columns list
the number of predicted ESE hexamers
in the affected region of the wild-type
and mutant sequences, respectively.
(C) Locations of all mutations listed in
(B) are indicated relative to the exon-
intron structure of the HPRT gene by
red arrows. Exon sizes are to scale;
intron sizes are not. Mutations that
alter RESCUE-predicted ESEs are shown
below the exon-intron schematic,
numbered according to (B). Mutations
labeled E3 X disrupt predicted ESE
hexamers; those labeled X3 E create predicted ESEs.

natural ESEs; the other five (5A/3G, 5B/3A, SELEX motif for SRp20 (table S2). The (17)]. For 14 of these cases, all predicted
5D, 3B, and 3H) are not similar to known, match between cluster 3C, consensus GA- ESE activity is lost in the mutant sequence,
mutationally defined ESEs. Many SELEX CRA, and the 9G8 binding SELEX motif as compared to only three positions where
motifs have not been confirmed by mutation- AGACKACGAY (K ⫽ G or T) is also mutations create predicted ESEs not
al analysis and so are not treated as “known” reasonable. These similarities suggest plau- present in the wild-type sequence. This ra-
enhancers here. On the other hand, some sible and testable trans-factor interactions. tio (14/3 ⫽ ⬃4.7) of predicted ESE-down
functional SELEX and all binding SELEX To evaluate the quality of RESCUE pre- mutations to predicted ESE-up mutations is
results identify trans-factor interactions, dictions in the context of a natural gene higher than that reported recently (43/21 ⫽
which are often not known for natural ESEs. locus, we considered the published set of ⬃2.0) for an approach using weight matri-
Databases of functional SELEX and binding exon mutations that cause exon skipping in ces based on functional SELEX results [see
SELEX consensus motifs were also con- the HPRT gene, mutation of which is asso- table 1 of (25)]. Thus, RESCUE-ESE should
structed by surveying the literature (tables ciated with Lesch-Nyhan syndrome (23, prove useful as a predictive tool for analyzing
S5 and S6, respectively) and were com- 24). The HPRT cDNA was considered as a the splicing phenotypes of exonic mutations
pared to RESCUE-ESE hexamers. All set of overlapping hexanucleotides, each or polymorphisms associated with human
“complete” matches spanning either the en- beginning one base 3⬘ of the beginning of disease. The general RESCUE approach
tire hexamer and/or the entire SELEX motif the previous hexamer. In this representa- could also be applied to identify enhancers or
are listed in table S1 (e.g., the hexamer tion, a point mutation disrupts six overlap- silencers of other processes such as polyade-
GAAGAA would be counted as a match to ping hexamers in the wild-type sequence, nylation or transcription.
GARGAR, YGARGAR, or ARGAR, but creating six new hexamers that are point
not to YARGAR). There were matches to mutants of the wild-type hexamers. Simi-
some of the clusters that lacked similarity larly, a one-base deletion at a given posi- References and Notes
to known natural ESEs, as well as several tion disrupts six wild-type hexamers and 1. L. P. Lim, C. B. Burge, Proc. Natl. Acad. Sci. U.S.A. 98,
11193 (2001).
matches to the purine-rich cluster 5C/3D. creates five new ones in their place. Of the 2. D. L. Black, RNA 1, 763 (1995).
The consensus sequence for cluster 5D, 30 HPRT mutations that alter hexamer 3. K. K. Nelson, M. R. Green, Genes Dev. 2, 319 (1988).
TC[GT]TC, is a particularly good match composition and cause exon skipping, 16 4. R. Reed, T. Maniatis, Cell 46, 681 (1986).
for the functional SELEX motif UCUUC positions coincide with at least one predict- 5. B. J. Blencowe, Trends Biochem. Sci. 25, 106 (2000).
6. L. Cartegni, S. L. Chew, A. R. Krainer, Nature Rev.
and also matches fairly well to positions 2 ed enhancer motif (Fig. 5). Overall, a total Genet. 3, 285 (2002).
to 6 of the SRp20-binding SELEX motif of 31 predicted ESE hexamers are present 7. L. R. Coulter, M. A. Landree, T. A. Cooper, Mol. Cell.
YWCUUCAU (Y ⫽ C or T; W ⫽ A or T). in positions overlapping the mutated posi- Biol. 17, 2143 (1997).
The cluster 3H consensus (ACT T) also tions in the wild-type sequence, whereas 8. H. X. Liu, M. Zhang, A. R. Krainer, Genes Dev. 12, 1998
(1998).
matches fairly well to two functional only seven predicted ESE hexamers are 9. H. X. Liu, S. L. Chew, L. Cartegni, M. Q. Zhang, A. R.
SELEX motifs and to another binding present in the mutant sequences [P ⬍ 0.01 Krainer, Mol. Cell. Biol. 20, 1063 (2000).

1012 9 AUGUST 2002 VOL 297 SCIENCE [Link]


REPORTS
10. T. D. Schaal, T. Maniatis, Mol. Cell. Biol. 19, 1705 19. D. L. Black, Genes Dev. 5, 389 (1991). tions. Supported by a Functional Genomics Innovation

㛬㛬㛬㛬
(1999). 20. Z. Dominski, R. Kole, Mol. Cell. Biol. 11, 6075 (1991). Award from the Burroughs Wellcome Fund (C.B.B.,
11. H. Tian, R. Kole, Mol. Cell. Biol. 15, 6291 (1995). 21. , Mol. Cell. Biol. 12, 2108 (1992). P.A.S.) and by NIH grant 1 R01 HG02439-01 (C.B.B.).
12. S. M. Berget, J. Biol. Chem. 270, 2411 (1995). 22. R.-F. Yeh, C. B. Burge, data not shown. Supporting Online Material
13. C. F. Bourgeois, M. Popielarz, G. Hildwein, J. Stevenin, 23. M. Tu, W. Tong, R. Perkins, C. R. Valentine, Mutat. Res. [Link]/cgi/content/full/1073774/DC1
Mol. Cell. Biol. 19, 7347 (1999). 432, 15 (2000). Materials and Methods
14. A. Kanopka, O. Muhlemann, G. Akusjarvi, Nature 381, 24. C. R. Valentine, Mutat. Res. 411, 87 (1998). Figs. S1 to S5
535 (1996). 25. H. X. Liu, L. Cartegni, M. Q. Zhang, A. R. Krainer, Tables S1 to S6
15. L. M. McNally, M. T. McNally, Mol. Cell. Biol. 18, 3103 Nature Genet. 27, 55 (2001). References
(1998). 26. We thank B. Blencowe, L. Chasin, L. Lim, D. Lipman, and
16. B. R. Graveley, RNA 6, 1197 (2000). D. Riordan for helpful comments on the manuscript; H. 9 May 2002; accepted 2 July 2002
17. Supporting data are on Science Online. Cargill for help with the figures; T. Cooper for Published online 11 July 2002;
18. J. D. Thompson, D. G. Higgins, T. J. Gibson, Nucleic generously providing us with the SXN minigene con- 10.1126/science.1073774
Acids Res. 22, 4673 (1994). struct; and anonymous reviewers for helpful sugges- Include this information when citing this paper.

Microbial Reefs in the Black Sea fate reduction (SR) and as an organic carbon
source for cell synthesis.

Fueled by Anaerobic Oxidation In the northwestern Black Sea, hundreds of


active gas seeps occur along the shelf edge west
of the Crimea peninsula at water depths be-
of Methane tween 35 and 800 m (14). At some of the
shallow Crimean seeps, microbial mats were
Walter Michaelis,1* Richard Seifert,1 Katja Nauhaus,2 found associated with isotopically light carbon-
Tina Treude,2 Volker Thiel,1 Martin Blumenberg,1 ates. Aspects of the microbiology, sedimentol-
Katrin Knittel,2 Armin Gieseke,2 Katharina Peterknecht,1 ogy, mineralogy, and selected biomarker prop-
erties of these deposits were recently described
Thomas Pape,1 Antje Boetius,3 Rudolf Amann,2 (11, 15–17). We explored the seeps on the
Bo Barker Jørgensen,2 Friedrich Widdel,2 Jörn Peckmann,4 lower Crimean shelf with the manned submers-
Nikolai V. Pimenov,5 Maksim B. Gulin6 ible JAGO from aboard the Russian R/V Pro-
fessor Logachev. During dives to a seep area at
Massive microbial mats covering up to 4-meter-high carbonate buildups prosper 44°46⬘N, 31°60⬘E, we discovered a reef con-
at methane seeps in anoxic waters of the northwestern Black Sea shelf. Strong sisting of up to 4-m-high and 1-m-wide micro-
13
C depletions indicate an incorporation of methane carbon into carbonates, bial structures projecting into permanently an-
bulk biomass, and specific lipids. The mats mainly consist of densely aggregated oxic bottom water at a depth around 230 m
archaea ( phylogenetic ANME-1 cluster) and sulfate-reducing bacteria (Desul- (Fig. 1A). These buildups are formed by up to
fosarcina/Desulfococcus group). If incubated in vitro, these mats perform 10-cm-thick microbial mats that are internally
anaerobic oxidation of methane coupled to sulfate reduction. Obviously, stabilized by carbonate precipitates (Fig. 1A).
anaerobic microbial consortia can generate both carbonate precipitation and From holes in these structures, streams of
substantial biomass accumulation, which has implications for our understand- gas bubbles emanate into the water column
ing of carbon cycling during earlier periods of Earth’s history. (Fig. 1A). The gas contains about 95% meth-
ane [see supporting online material (SOM)].
Until recently, it was believed that only aerobic existence of methane-consuming associations In cross section, the outside of the soft mat
bacteria, depending on oxygen as an electron of archaea and sulfate-reducing bacteria (SRB) has a dark gray to black color (Fig. 1B).
acceptor, build up substantial biomass from in anoxic marine sediments (3, 4). Inside the structure, most of the mat is pink to
methane carbon in natural habitats (1). Because Microorganisms capable of anaerobic brownish. The interior rigid parts are porous
biogenic methane is strongly depleted in 13C, a growth on methane have not been cultivated carbonates (aragonite and calcite with up to
worldwide negative excursion in the isotopic so far, and the biochemical pathway of the 14% MgCO3). Much of the structures con-
signature of organic matter around 2.7 Ga (1 anaerobic oxidation of methane (AOM) re- sists of interconnected, irregularly distributed
Ga ⫽ 109 years) ago was taken as an argument mains speculative. Analyses of depth profiles cavities and channels filled with seawater and
for methanotrophy and, consequently, an early and radiotracer studies in marine sediments gases. Apparently, the cavernous structure of
oxygenation of the Earth’s atmosphere (2). (5, 6) as well as molecular (7–10) and petro- these precipitates enables methane and sul-
However, recent investigations have shown the graphic studies (11) argue for AOM (12) as a fate to be transported and distributed through-
crucial process that channels 13C-depleted out the massive mats. Smaller microbial
methane carbon into carbonate and microbial structures and nodules from nearby areas
1
Institute of Biogeochemistry and Marine Chemistry,
University of Hamburg, Bundesstrasse 55, 20146 biomass. Through AOM mediated by consor- were of the same morphology, with compact
Hamburg, Germany. 2Max Planck Institute for Marine tia of archaea and SRB, methane is oxidized mat enclosing calcified parts and cavities.
Microbiology, Celsiusstrasse 1, 28359 Bremen, Ger- with equimolar amounts of sulfate, yielding Obviously, the microorganisms do not grow
many. 3Alfred Wegener Institute for Polar and Marine carbonate and sulfide, respectively (13). Gen- on preformed carbonates but induce and
Research, 27515 Bremerhaven, and International Uni-
versity Bremen, 28725 Bremen, Germany. 4Geowis- eration of alkalinity favors the precipitation shape their formation. Stable carbon isotope
senschaftliches Zentrum, University of Göttingen, of methane-derived bicarbonate according to analyses of the carbonates yielded ␦13C val-
Goldschmidtstrasse 3, 37077 Göttingen, Germany. the following net reaction: ues ranging from ⫺25.5 to ⫺32.2 per mil
5
Institute of Microbiology, Russian Academy of Sci- (‰) [for methods, see (11)]. Compared with
ences, pr. 60-letiya Oktyabrya 7, k. 2, Moscow, CH4⫹SO42⫺⫹Ca2⫹3 CaCO3⫹H2S⫹H2O
117811, Russia. 6Institute of Biology of Southern
the ␦13C values of dissolved inorganic carbon
Seas, National Academy of Sciences of Ukraine, pr. Here, we provide evidence that vast amounts in the Black Sea water column from ⫹0.8‰
Nakhimova 2, Sevastopol, Ukraine. of microbial biomass may accumulate in an at surface to ⫺6.3‰ at depth (18), these
*To whom correspondence should be addressed. E- anoxic marine environment because of the values indicate that a major portion of the
mail: michaelis@[Link] use of methane as an electron donor for sul- carbonate originates from the oxidation of

[Link] SCIENCE VOL 297 9 AUGUST 2002 1013

You might also like