0% found this document useful (0 votes)
1 views15 pages

Tutorial Genomics ZL

This document is a tutorial on genomics and gene prediction using the Ensembl genome browser and various bioinformatics tools. It covers the basic functionalities of Ensembl, how to use BioMart for gene data analysis, and methods for predicting open reading frames and tRNA genes. The tutorial includes practical exercises and examples to enhance understanding of gene analysis and comparative genomics.

Uploaded by

nadianparisa1999
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
1 views15 pages

Tutorial Genomics ZL

This document is a tutorial on genomics and gene prediction using the Ensembl genome browser and various bioinformatics tools. It covers the basic functionalities of Ensembl, how to use BioMart for gene data analysis, and methods for predicting open reading frames and tRNA genes. The tutorial includes practical exercises and examples to enhance understanding of gene analysis and comparative genomics.

Uploaded by

nadianparisa1999
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Genomics and gene

prediction
concepts and methods
Bioinformatics

Lecturer: Zelmina Lubovac

[Link]@[Link]
Contents

1 Introduction.......................................................................................................................................... 3

1.1 Basic functions of Ensemble .......................................................................................................... 3

1.2 Ensembl BioMart ........................................................................................................................... 9

1.3 Gene prediction (this is for the last lab, after lecture 8) ............................................................. 11

2
1 Introduction
In this tutorial, you will explore the basic functionalities of Ensembl, a comprehensive
genome browser that provides detailed information on genes, their sequences, and related
genomic data. The tutorial is designed to give you hands-on experience with gene search,
transcript analysis, protein domains, synteny comparisons, and gene prediction tools such as
BioMart, tRNAscan-SE, and ORF Finder.

By the end of this tutorial, you will be able to:

• Navigate Ensembl to search for genes and explore their genomic contexts.

• Use BioMart to filter, export, and analyze gene and transcript data.

• Predict open reading frames (ORFs) and tRNA genes using tools from external
databases.

• Compare genomes across species and understand the significance of conserved


regions.

The tutorial is divided into three parts:

Part 1: Basic functions of Ensembl, including navigating the genome browser, exploring
transcripts, protein domains, and synteny.

Part 2: Using the Ensembl BioMart tool to retrieve and manipulate gene data.

Part 3: Gene prediction using ORF Finder and tRNAscan-SE, alongside practical exploration of
bacterial genomes.

Each part will guide you through specific tasks designed to deepen your understanding of
genomics and gene analysis tools.

1.1 Basic functions of Ensemble


Open your web browser and navigate to the Ensembl website
([Link] Search for TP53 gene, by entering it in the search
box (see Fig 1) and select human genome (Homo Sapiens) from the search result.

3
Figure 1: Ensambl interface

When you click on Go and select the top hit on the resulting list, you will you will be directed
to the gene summary page, as shown in the Fig 2. This Gene tab contains all the information
associated at the gene level such as: location on the chromosome, transcripts etc. At the top
of the gene summary, the number of transcripts, or splice variants, are shown in a table.

4
Figure 2: Gene tab provides an overview of gene transcripts.

Exercises Ensembl

1. Which chromosome is gene TP53 located on?

Answer: TP53 is located on the chromosome 17.

2. How many transcripts does it have and how many of those are protein coding?

Answer: TP53 has 33 transcripts, 28 of which are protein coding.

3. Select the longest transcript and answer following questions:

a. How many exons are shown for this transcript?

Answer: Sort the transcript table based on the length (bp). The ID of the
longest transcript is ENST00000714409.1. It has 10 exons, of which 9 are
coding.

b. What is the length of the transcript and the corresponding amino acid
sequence?

Answer: This transcript contains of 3 430 nucleotide bases and 367 amino
acid residues.

5
c. Select the fifth transcript ID in the sorted table and check its number of exons.
Does it differ from the first transcript?

Answer: The fifth transcript (ENST00000610292.4) has 10 exons, of which 8


are coding, 2639 nt bases and 354 aa residues.

4. Click the protein link belonging to the longest transcript (in the “Protein” column of
the table) and inspect the "Protein summary" that will appear below the transcript
table.

a. What protein domains are shown?

Answer: Domains include “p53-like transcription factor, DNA-binding domain


superfamily” (SuperFamily), “p53 tumor suppressor family” (Prints), “Cellular
tumor antigen 53, transactivation domain 2” (Pfam), “p53, DNA-binding
domain” (Pfam), “p53, tetramerisation domain”m (Pfam), etc (see Fig 3).

Figure 3: Overview of protein domains.

6
b. Click on the p53 transactivation domain and view its Pfam, entry. What can you
find out about the function of the domain? What is the main role (very briefly
summarized) of p53 transcription factors in cancer?

Answer: In the Pfam entry PF08563 you can read about how the p53
transactivation domain is involved in regulating the activation of p53
transcription.

5. Search for the chromosome region where TP53 is located by entering


17:7600000‐7750000 in the search field. (Can alternatively be written as
17:7,600,000‐7,775,000). Note that the menu to the left changes when you look at a
region. Go to "Region overview" in the menu and inspect the genome browser.

a. What other genes are located near TP53? Write down the names of their
proteins and save this list for later use (in other questions).

Answer: There are 7 genes, in addition to TP53: FXR2, SHBG, SAT2, ATP1B2,
WRAP53, EFN53, DNAH2 (see Fig 4).

Figure 4: Region overview for TP53 and closely located genes.

6. Go to the “Synteny” section in the left menu. You should see a comparison of human
chromosome 17 with a chromosome from another species.

a. Which species? On which chromosome would we expect to find TP53 in that


species?

Answer: Mouse chromosome 11 (see Fig 5).

7
Figure 5: Synteny between Human chr. 17 and mouse.

b. Change the comparison species to rabbit and check the same information.

Answer: If we choose rabbit, the synteny is to rabbit chromosome 9 (Fig 6).

Figure 6: Synteny between Human chr. 17 and rabbit.

7. Go back to the genome browser display of the region 17:7600000‐7750000 in the human
genome. In the lower part of the menu to the left, click on the link “Export data” (see Fig 7).
Download the nucleotide sequence of this region (150 kbp) and save it in a text file. Give the
file an appropriate name and save it in your working directory. It will be used in a later
exercise.

8
Figure 7: Exporting FASTA sequence for specific region.

1.2 Ensembl BioMart


Exercises Ensambl BioMart

8. Use the BioMart to find all genes occurring in the region 17:7600000‐7750000 of the
human genome. Make sure that you get only genes, no transcripts (i.e. only ENSG
entries). Use Filters in the menu to the left to specify chromosome, and the region.
Click on Attributes and choose to display Gene stable IT and Gene name. You need to
click on Count on the top of the menu to get the number of genes in this region.
Finally, to get the table with the genes, click on Results. Your resulting table should
look like the one in Fig 8 bellow.

a. How many genes do you find, and does the number correspond to what you
saw in the genome browser? Verify that TP53 is among the genes found.

Answer: There are 8 genes, and TP53 is one of them. The number of genes
corresponds to the number we saw in the genome browser.

9
Figure 8: Resulting table with TP53 and 7 closely located genes on chr. 17.

b. Change the “Attributes” settings so that you also get transcript entries. How
many transcripts do you get for the first gene in the result table?

Answer: There are 8 transcripts for the first gene, FXR2 (see Fig. 9).

Figure 9: Resulting table with transcripts for FXR2.

9. Change the settings “Filters” and “Attributes” settings so that you find answers to the
following questions:

a. How many genes in total on the same chromosome as TP53? (Note that you
don’t need to download them, it is sufficient to use the “Count” button).

Answer: First, change settings in Attributes by removing “Transcripts stable


ID” since you are only looking for genes. Second, remove the ”Coordinates”
settings in Filters since you are looking at the whole chromosome 17. By
clicking on ”Count”, we found that there are a total of 3191 genes on
chromosome 17 (see Fig 10).

10
Figure 10: Resulting count for genes on Human chromosome 17.

b. Of the genes on this chromosome, how many have SWISSPROT accession


numbers?

Answer: There are 1118 genes on chr 17 which have SwissProt accession
numbers. To apply this filter you go to ”Filters”, then ”GENE”, then ”Limit to
genes (external references)”, then scroll the list until you find ”With
UniProtKB/SwissProt ID(s)”.

Figure 11: Resulting count for genes with SwissProt ID on chr. 17.

1.3 Gene prediction (this is for the last lab, after lecture 8)
Now we are going to explore the ORF finder (Open Reading Frame Finder) tool at the NCBI
web server. ORF Finder is used to identify open reading frames in a DNA sequence, which
are regions that have the potential to code for proteins.

11
10. Navigate to the following web page: [Link] and
use the sequence you downloaded in question 7, to find open reading frames in
region 17:7600000‐7750000.

a. How many bases is the longest ORF found? Was it found in a forward or a
reverse orientation frame?

Answer: By using default settings (>75 bp), you get 323 ORFs. The longest one
is 561 bp and is found on forward strand (+) (see Fig 12).

Figure 12. Results from Open Reading Frame finder.

b. Adjust the minimum ORF length to 300 and observe what happens with the
number of predictions.

Answer: Adjustment to >300 bp will result in much fewer ORFs, 16 in total.


For parameter settings, see Fig 13.

12
Figure 13. Parameter setting in Open Reading Frame finder.

11. Go to Ensembl Bacteria and look at the first 20,000 bp of the chromosome of
Mycoplasma genitalium strain G37. According to the view in the genome browser,
how many protein coding genes are there in this region? Download the sequence (of
20 kbp) to a text file in Fasta format.

Figure 14: Ensambl Becteria genome browser view for region in Mycoplasma
genitalium strain G37

Answer: The genome browser shows 16 protein coding gens (red color) and one RNA
gene (in purple), as shown in Fig 14.

12. Use tRNAscan‐SE to predict tRNA genes in the 20 kbp sequence fragment from
Mycoplasma. How many tRNA genes are predicted by this tool?

Navigate to this page: [Link] Choose the type of organism

13
from the drop-down menu. In this case, it is set to Bacterial. Keep default settings. You can
paste your sequence directly into the query box (as shown in Fig 15) in FASTA format. Click
the "Run tRNAscan-SE" button to start the tRNA search on the input sequence.

Figure 15: tRNAscan-SE Input Interface.

Answer: tRNAscan-SE predicts two tRNA genes, at the positions shown below in the Results
table (as shown in Fig 16).

Figure 16: Resulting table from tRNAscan-SE

13. Use ORF finder to predict open reading frames in the same sequence.

a. How many ORFs of length 300 bp and longer do you find?

Answer: ORF-finder finds 16 open reading frames of length >= 300 nt.

b. Does the number correspond approximately to the number of known genes?

Answer: The number of ORF:s corresponds exactly with the number of genes
according to Ensembl.

c. Submit the longest ORF to blastp and check if the top hit corresponds to one of
the genes shown in the Ensembl genome browser

14
Answer: The longest ORF has location and length corresponding to the gyrA
gene in Ensembl, starting at 4,812 and ending at 7,322 (see Fig 17). When
running a Blast search with the longest ORF as query sequence, the first hit
that comes up is gyrase subunit A, i.e. gyrA (see Fig 18).

Figure 17: Results from ORF-finder compared to Ensambl genes.

Figure 18: Results from Blast search with the longest ORF

15

You might also like