0% found this document useful (0 votes)
19 views11 pages

UCSC Genome Browser Tutorial Guide

Uploaded by

Amir
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views11 pages

UCSC Genome Browser Tutorial Guide

Uploaded by

Amir
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Tutorial 1

Introduction to the UCSC Genome Browser

Try these exercises, following the step-by-step instructions below:

1) Explore the browser, learning how to turn on and off the tracks.

2) Retrieve how much, and in which cells, a gene is expressed.

3) Identify expression regulation signals in a cell line.

Then try yourself:

4) Look for the PPP1R1B gene (protein phosphatase 1, regulatory (inhibitor)


subunit 1B) in the human genome, release Feb. 2009. Check the number of
splicing variants reported by UCSC, Ensembl and RefSeq. On which strand is
the gene transcribed? Which is the closest gene to PPP1R1B?

5) In which cells/tissues is the gene PPP1R1B more expressed? Are gene


expression microarray data in agreement with those from the GTEx project?

6) Collect some information on the expression regulation of PPP1R1B


1) Explore the browser, learning how to turn on and off the tracks.

1 Connect to the UCSC Genome Browser homepage, [Link].


2 Access the Gateway, moving your mouse over the Genome Browser link in the blue
bar on top of the screen, and click on Reset All User Settings.

If you are asked which mirror you want to access, select Let me stay here.
3 Now you are in the Gateway, from which you can have access to the different
available genomes. You need to point the browser to one of the available genome
assemblies of the human genome. In the Human Assembly field, select Feb. 2009
(GRCh37/hg19). Below, in the Position/Search Term field, in Current Position,
enter the following coordinates on chromosome 21 by copying and pasting the
following: chr21:33,031,597-33,041,570. Click on the Go button to direct the browser
to these genomic coordinates.

4 The browser is now displaying, in the Genome Viewer panel, a genomic region
containing the SOD1 gene. Also, the default tracks are displayed, including ChIP-Seq
data, DNase I cut sensitive sites, evolutionary conservation, histone H3 acetylation,
and others. Let’s clear the panel by deactivating all these tracks. Just below the
Genome Viewer panel you can find a series of buttons, among which the hide all
button. Click on it. Doing that, the Genome Viewer will display just the genomic
coordinates of the region that you are watching.

5 Now let’s have a look at the SOD1 gene structure, its functions and how different
databases might report different gene structures for it. Scrolling down the page, you
can find the track control panel, which is organized in different groups. Look for the
Genes and Gene Predictions group. If the available tracks within this groups are
not shown, click on the + button to the left of the track group name to expand it. The
first track is UCSC Genes, reporting genes as defined by the UCSC pipelines. A drop
down menu allows you to select the track display mode. The track is not displayed
now, hence its display mode is hide. From the UCSC Genes drop down menu select
pack, then click on any refresh button on the right.
Now, the Genome Viewer displays the SOD1 exons (as black boxes) and introns:

How many exons this gene contains? If you hover the mouse over an exon or an
intron, the numbering of it is displayed. Clicking on the gene name on the left of its
track in the Genome Viewer, you can access the gene entry, reporting links to its
nucleotide and protein sequence, its expression in different tissues and cell types,
the encoded protein three-dimensional structure (if known), its role in pathologies, its
role in signaling or metabolic pathways, etc. Have a look at all the information
provided. Then, go back to the previous page by using the Back button of the
internet browser (Firefox, Chrome, etc.).
6 Let’s zoom in to see the nucleotide sequence of the SOD1 second exon. On top of
the screen, you can see a series of zoom in buttons. Click on 10x.

The second exon of SOD1 should now be towards the left side of the panel. Let’s
move it to the screen center by clicking the panel with the mouse left button, and
dragging it to the right while keeping the button pressed.

When the exon is centered in the viewer, click again on zoom in 10x. You should
now see the exon nucleotide sequence, and the amino acids encoded by the exon.

In the track group Mapping and Sequencing (if the track of this group are not show,
expand it using the + button), activate the Base Position track by selecting full
mode from the drop down menu, then click refresh. Now you should see the
sequence translation in all possible reading frames, one of which coincides with the
reading frame of the SOD1 exon. The translation of the genomic sequence in all
possible reading frames can be used to scan the genome looking for all possible
open reading frames (ORFs), which might indicate unknown genes.
7 Let’s now see how different databases report the SOD1 gene structure. Click twice
on zoom out 10x. If needed, center the SOD1 gene in the panel by clicking and
dragging the screen with the mouse left button. In the track group Genes and Gene
Predictions, activate the tracks Ensembl Genes and NCBI RefSeq, both in dense
mode, then refresh.

In the Genome Viewer two new tracks are now displayed, one describing SOD1 as
reported by the RefSeq database (it’s the one colored in blue), the other reporting the
gene as reported by the Ensembl database (in red). You can see that RefSeq reports
the same exon and intron structure as UCSC, while the Ensembl track shows some
differences. For example, the first exon reported by Ensembl is longer, and there are
few additional exons.

8 Let’s also see how SOD1 splicing is reported by these databases. In the track group
Genes and Gene Predictions, change the viewing mode of the UCSC Genes,
Ensembl Genes and NCBI RefSeq tracks, by selecting full this time, then refresh.
In the Genome Viewer, you should now see that both RefSeq and UCSC reports
only one splicing variant for SOD1, while Ensembl reports 4 variants. Moreover,
Ensembl reports two additional transcribed exons in the same region (shown as red
boxes), just below the 4 SOD1 splicing variants, that are not linked to SOD1 exons.
Click on the first one (having identifier ENST00000458922).

The next page is a description of this transcript, which is a snoRNA, a small nucleolar
RNA. You can read that it is transcribed from the + strand, the same of SOD1, and
that it is found into a SOD1 intron. This is common for snoRNA, which often derive
from maturation of introns removed by splicing. Now go back to the previous page
and click on the other transcribed region below the SOD1 splicing variants
(ENST00000609934). In the next page, you can find out that this is an antisense
RNA, transcribed from the – strand, that could bind SOD1 transcripts preventing their
translation.
9 Now let’s get information on the SOD1 encoded protein. In the Genes and Gene
Predictions tracks group, activate the UniProt track choosing pack from the drop
down menu, then click on refresh. In the Genome Viewer several new tracks will
appear, describing different features of the encoded protein, including post-
translational modifications (phosphorylation, acetylation, etc.), known mutations,
disulfide bonds, domains, the secondary structure, binding sites, etc.
2) Retrieve how much and in which cells a gene is expressed.
1 Let’s start from the beginning. From the UCSC Genome Browser homepage,
[Link], access to the Gateway, moving the mouse cursor on the link
Genome Browser on top of the screen and clicking on Reset All User Settings.
Verify that the browser is connecting to the human genome. In Human Assembly
select Feb. 2009 (GRCh37/hg19). In the text box Position/Search Term, in Current
Position, enter the following coordinates on chromosome 21: chr21:33,031,597-
33,041,570. Click the Go button to access these genomic coordinates.
2 The browser is now showing again the region containing the SOD1 gene, and the
default tracks. Expand the visual by clicking 3x in zoom out. You should see the
beginning of the SCAF4 gene, downstream to SOD1. Move SCAF4 to the center of
the screen by clicking >>> in move (you can find it on the top of the screen, just over
the viewer), or by dragging the screen with the mouse.

If the whole SCAF4 gene does not fit in the screen (it is very long), use zoom out
1.5x until you can see the whole gene. Notice how 5 different SCAF4 splicing variants
are shown in the UCSC genes track, while the NCBI RefSeq track reports only one
variant. Now look at the track labeled Layered H3K27Ac. This track shows the
results of ChIP-seq experiments targeting Histones 3 acetylated on lysine 27, a signal
usually associated to active transcription, and often found in the promoters of
transcribed genes. You should see peaks indicating the presence of this epigenetic
signal on the right side of SCAF4.

Why are these signals found on the right side of the gene, while for example they are
found on the left of SOD1?
3 Click on the hide all button below the viewer to deactivate all the active tracks, to
clean the screen.
4 Let’s turn on the UCSC Genes track in pack mode, and the GTEx gene V8 track in
the Expression group, also in pack mode, then refresh. Then zoom out 3x. You will
see in the viewer a series of histograms, one for each gene shown in the viewer,
reporting gene expression in different human tissues as measured with RNA-Seq in
the context of the GTEx project. If you followed the instructions up to this point, you
should see expression histograms for the SCAF4, SOD1, SNORA81 (which is a
snoRNA, a small nucleolar RNA) genes, e for a few other genes, including antisense
RNAs (shown as AP followed by a numeric code). Each histogram bar show the
expression level in a different tissue.
3) Identify expression regulation signals in a cell line.

1 Let’s start again from the beginning. From the UCSC Genome Browser homepage,
[Link], access to the Gateway, moving the mouse cursor on the link
Genome Browser on top of the screen and clicking on Reset All User Settings.
Verify that the browser is connecting to the human genome. In Human Assembly
select Feb. 2009 (GRCh37/hg19). In the text box Position/Search Term, in Current
Position, you should still see the previous coordinates in chromosome 21
(chr21:33,031,597-33,041,570). If this is not the case, copy & paste these
coordinates. Click the Go button to access these genomic coordinates.
2 Expand the view by clicking 10x in zoom out. Click on hide all in the middle of the
screen just below the viewer to deactivate all tracks
3 Scroll down until you find the group of tracks Genes and Gene Prediction, and
activate the track UCSC Genes in pack mode. Then, look for the track group
Regulation, and activate the track CpG Islands as show, then refresh.

4 CpG islands are shown as green boxes, with a label indicating the number of CpG
dimers found within the island. Notice that there is a CpG island that overlaps with
the promoter and first exon of the SOD1 gene.

Click on the green box to get information on the island (its length, how many CpG
dimers it contains, etc.). Then, go back by using the Back button of the internet
browser
5 In the group Regulation, activate the track ENC DNA Methyl as show, then
refresh. This track shows the methylation state as measured in different cell lines.
Look for methylation sites (depicted as vertical bars) detected in the GM12878 cell
line (their name starts with GM12878, on the right side of the Viewer), you will see
that many sites are colored in green. Click on one of these green bars to get a
description of the meaning of the color-coding. Can you say whether the CpG island
is methylated or not in this cell line? Based on its methylation, can we expect that the
SOD1 gene is expressed in this cell line? Remember that C methylation is generally
a transcription repression signal.
6 In the track group Regulation, activate the track ENCODE Regulation as show, and
then refresh. The ENCODE Regulation track show a series of epigenetic signals
such as the acetylation of lysine 27 of histone H3 (H3K27Ac), the DNase I cut
sensitivity (DNase clusters), which indicates accessible chromatin regions, and
transcription factors binding sites determined by ChIP-Seq (Txn Factor ChIP).
7 In the Viewer, click on the histone modification track H3K27Ac to access to its
description and configuration page, as shown in the figure:

Have a look at the description of this track. The track reports the ChIP-Seq peaks
detected in seven different cell lines. Unselect all cell lines, except the GM12878 line.

Set Display mode as full (at the top of the page), then Submit (on top of the page,
near the Display mode control). The browser will show again the Viewer, but this
time showing only the selected cell line as an orange histogram. You should see a
pronounced peak in the SOD1 promoter. Is this evidence suggesting that SOD1 is
expressed or not in the GM12878 cell line? Is this signal in agreement with the
methylation state of the CpG island?
8 Expand the track Txn Factor ChIP, which is part of the track ENCODE Regulation
track, by clicking on it in the Viewer. You are going to see a set of black or grey
boxes or bars, each one corresponding to the binding site for a transcription factor or
other DNA-binding proteins, detected by immunoprecipitation. Look for the black box
labelled POLR2A, in the SOD1 promoter, and click on it.
A description page will open, reporting in a table the cell lines in which this protein
was detected binding at that locus. The second column of the table reports a score
(signal), which is higher for more intense signals, indicating strong binding. The
fourth column (cellType) indicates the cell type. Look for GM12878 lines; can you
conclude that POLR2A is strongly binding to SOD1 promoter in these cells? Click on
POLR2A in the column factor. Which protein it is? Is the POLR2A binding
suggesting that SOD1 is expressed, or not?
9 Let’s now verify whether SOD1 is expressed in GM12878 cells. Use the track GNF
Atlas 2 in the group Expression. Configure the track as you did in Exercise 2.
GM12878 is a lymphoblast cell line. GM12878 is not included in this track, but you
can look for the expression in a similar cell line (721 B lymphoblast). Is the detected
expression consistent with the methylation and ChIP signals?

You might also like