100% found this document useful (1 vote)
106 views14 pages

Bioinformatics Essay Assignment Guide

This document provides instructions for a practical assignment on analyzing a DNA sequence using bioinformatics tools. It involves three main tasks: 1. Extracting the E. coli K12 recA gene sequence with accession number V00328 from the European Nucleotide Archive database. 2. Analyzing the extracted sequence to find open reading frames, suitable restriction enzyme sites for cloning, and primers for PCR amplification of the gene region. Tools used include NCBI ORF Finder and primer design tools. 3. Performing BLAST searches and PSI-BLAST iterations to find structurally and sequentially similar proteins like tyrosine tRNA ligase and tryptophan tRNA ligase, and determine evolutionary relationships

Uploaded by

visini
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
100% found this document useful (1 vote)
106 views14 pages

Bioinformatics Essay Assignment Guide

This document provides instructions for a practical assignment on analyzing a DNA sequence using bioinformatics tools. It involves three main tasks: 1. Extracting the E. coli K12 recA gene sequence with accession number V00328 from the European Nucleotide Archive database. 2. Analyzing the extracted sequence to find open reading frames, suitable restriction enzyme sites for cloning, and primers for PCR amplification of the gene region. Tools used include NCBI ORF Finder and primer design tools. 3. Performing BLAST searches and PSI-BLAST iterations to find structurally and sequentially similar proteins like tyrosine tRNA ligase and tryptophan tRNA ligase, and determine evolutionary relationships

Uploaded by

visini
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Assignment Cover Sheet

Qualification Module Number and Title

HND in Biomedical Science BMS4009


Biomedical Techniques and Bioinformatics
Student name & No Assessor
Ms. Nilusha Weerasekra

Hand out date Submission data

Assessment type: Duration / Length Weighting of assignment


Essay of Assessment 20%
Types
1000 words

Learning declaration

I, …………………..., certify that the work submitted for this assignment in my own and
research sources are fully acknowledged.

Marks Awarded

First assessor
Second assessor
Agreed grade
Signature of the assessor Date

Signature of the assessor Date

Feedback Form
INTERNATIONAL COLLEGE OF BUSINESS AND TECHNOLOGY
Module Name : Biomedical Techniques and Bioinformatics
Student :
Assessor 1 : Ms. Nilusha Weerasekara
Assessor 2 :
Assignment : Practical Assignment on Bioinformatics

Areas for improvement:

Strong features of your work:

Marks Awarded:

LO : 6. Be able to understand the principles behind bioinformatics and independent to


carryout different types of bioinformatics related analyses using various tools and software.
Learning outcomes covered

Scenario and the Tasks


Bioinformatics is conceptualizing biology in terms of macromolecules (in the sense of physical-
chemistry) and then applying "informatics" techniques (derived from disciplines such as applied
math’s, computer science, and statistics) to understand and organize the information associated
with these molecules, on a large-scale. This assignment is designed to assess the principles
behind bioinformatics and independent to carryout different types of bioinformatics related
analyses using various tools and software.
Your report should cover all tasks listed below with a minimum 1000 and maximum 1500 word
count

Task 01 (LO 06)

1. Mentioned the applications of bioinformatics (5 Marks)

2. List the software and tools that can use in bioinformatics (5 Marks)

Task 02 (LO 06)

Genome Workbench offers researchers a rich set of integrated tools for studying and analyzing
genetic data. Users can explore and compare data from multiple sources including the NCBI
databases or the user’s own private data. Data analysis in Genome Workbench is supported by an
advanced suite of industry standard alignment tools including BLAST, Clustal, Kalign, MAAFT
among others. Learning and Discovery is accomplished along the way. When a researcher
studies and analyzes results with Genome Workbench, he or she views the data in novel ways
that leads to new understanding and discovery. Users are invited to take advantage of the
flexibility included via these tools to create phylogenetic trees, alignments, tabular view etc. of
data to graphically display data analyses in publications and presentations.

Using Blast software perform the following activities and submit your results
([Link]

Protein Bioinformatics resources at NCBI


Go to the NCBI website [Link] by entering the URL in the address field of
your browser.

Sequence Search

Protein serch

Task 1: Detecting Frameshift mutations Task 1: instructions 1. Search the sample sequence given
below against the protein database (nr) using the BLASTX program. Answer the questions that
follow.

sample_sequence
AGAAGAAGACATAGTAATTAGATCTGAAAATTTTACGAACAATGCTAAAACCATAA
TAGTACAGCTGAAG
GAATCTATAAAAATTAATTGTACAAGACCCAACAACAATACAAGAAAAAGTATACC
TATAGCAACGGGGG
GAGCAATTTATGCAACAGGAGACATAATAGGAGATATAAGACAAGCACATTGTAAC
CTTAGTAGAGACCA
ATGGGATAACACTTTAAGCCAGCTAGTTACAAAACTAAGAGAACAATTTGGGAATA
AACAATAGCCTTTAATCAATCCTCAGGAGGGGACCCAGAAATTGTAATGCACAGTTT
TAATTGTGGAGGAGAATTTTTCTACT
GTAATACAACACAGCTGTTTAATAGTACTTGGCCAACTAATAAAAAGTCTACTAACA
AAACAGGAAC
TATCACACTCCCGTGCAGAATAAAACAAATTATAAACAGGTGGCAAGAAGTAGGAA
AAGCAATGTATGCC
CCTCCCATCAAGGGACAAATTAGATGTTCATCAAATATTACAGGGATATTCTTAACA
AGAGATGGTGGTA
ACGCAAGCGATGAGACCGAGACCTTCAGACCTGGAGGAGGAAATA

2. For database hit AAL71600.1 which frame of the query sequence does alignment begin in? +1

3. At which nucleotide of the query sequence does the frame change?


TIAFNQSSGGDPEIVMHSFNCGGEFFYCNTTQLFNSTWPTN +STNKT TITLPCRIKQ
3. Using what you learnt about navigating through the NCBI database, find the nucleotide
sequence corresponding to the protein with the accession AAL71600.1. Describe the
steps you took to find it.

Go to NCBI, Select BLAST, Select BlastX, Copy and paste the sequence , BLAST, got
the list of hits, select the given hit, identified the frame sequence which had the mutation

4. Make a local alignment of the nucleotide sequence from 3. with the sample sequence.
Download your alignment and paste it in your assignment report. >AAL71600.1 envelope
glycoprotein, partial [Human immunodeficiency virus 1]

EEDIVIRSENFTNNAKTIIVQLKESIKINCTRPNNNTRKSIPIATGGAIYATGDIIGDIRQAH
CNLSRDQWDNTLSQLVT

KLREQFGNKTIAFNQSSGGDPEIVMHSFNCGGEFFYCNTTQLFNSTWPTNNTRSTNKTRT
ITLPCRIKQIINRWQEVGKA

MYAPPIKGQIRCSSNITGIFLTRDGGNASDETETFRPGGGN GGNASDETE

TFRPGGGN

5. Highlight the nucleotide that is bringing about the frameshift in the alignment you have
pasted in 4. NO frameshift

7. What other differences are there between the two nucleotide sequences? Do they also cause a
frameshift? Do they change the amino acid? First sequence had a 1+ frame shift but the 2 nd
sequence had no frame shift

Task 2.2 : Finding structurally similar proteins using PSI-BLAST

Tyrosine tRNA ligase (TyrRS) and Tryptophan tRNA ligase (TyrRS) are structurally similar
(Refs: 1, 2). Given structural similarity you would expect to find sequence similarity. However,
TyrRS and TrpRS share 13% sequence identity. We will use PSI-BLAST to find the sequence of
TyrRS.

1. Using a sequence database of your choice retrieve the protein sequence for E. coli
Tyrosine tRNA ligase (TyrRS), alias Tyrosyl-tRNA synthetase (sp|P0AGJ9.2).

mntvqeqmav irrgaveilv eaeleekike siakgvplri kagfdptapd lhlghtvliq


61 klkqfqelgh evcfligdft gmigdptgkn etrkpltreq vlanaqtyre qvfkildpak
121 tkvvfnsswm gpmtaadlig laarytvarm lerddfhkrf sgqqpiaihe flyplvqgyd
181 svalkadvel ggtdqkfnll vgrelqkqeg qrpqsvltmp llegldgvnk mskslgnyig
241 iteppreiyg kvmsisdelm iryyellsdv dlaglqqvkd gvagkasgah pmeskkalar
301 elvtrfhgqd qasqaemdfi qqfkqkeipd dipsvrmssd gpvwicrllv daglvasnge
361 arrmvkqggi kldgekivda dlevqpqgef vlqagkrrfa ritfgs

2. Open the BLASTP webpage on NCBI and paste the sequence you found in 1.

3. Submit a PSI-BLAST search against the SwissProt database, narrowing your search organism
to Bacteria. Keep the PSI-BLAST threshold at 0.005.

4. In your search results the description section will be split in two tables. The top table will
contain alignments with E-values below the cut-off and the lower table will show the hits above
the E-value cut-off.

a. What range of the % identities do you see in the alignments in both the tab. Do you get hits to
proteins other than TyrRS in the first table?

Table 1: Range is 23.98%. To 82. 67%

Table 2: Range is 30.13% to 27.54%.

c. Take a look at the Taxonomy report for this search. You will find a link to it just above the
Graphic Summary section. Which organism has the lowest alignment score? . Synechococcus
sp. PCC 7002

5. Run the 2nd iteration of PSI-BLAST using hits in the first table. This may take a while to run
so please be patient. Do you get hits to TrpRS on the 2nd iteration? Yes Using the Taxonomy
report for the 2nd iteration determine:
a. if you get a hit to TrpRS of E. coli. YES

b. which is the evolutionarily closest bacterial species to E. coli you get a TrpRS
. . . Escherichia coli UTI89 hit for? Hint: look at the number of hits in the lineage report.

c. how long is the TrpRS protein for this species? bles?

6. Run a 3rd iteration this time including the TrpRS hits.

a. Do you get a hit to TrpRS of E. coli?

b. If yes, what is the accession number for it?

c. What % identity does the query sequence (E. coli TrpRS) align with it (E. coli TyrRS)?

Task 3
DNA sequence analysis Practical

Introduction

DNA sequences can be extracted from public databases, or users may have their own piece of
DNA which they have sequenced, and which requires some basic analysis. In this practical we
explore basic analysis of a DNA sequence extracted from the database. Analysis of DNA
sequences usually entails finding important features encoded on the sequence. In this practical
we will search for open reading frames, try to find suitable restriction enzyme sites for cloning
most of the ORF, and design primers for PCR amplification of the ORF region
Tools used in this session

For extraction of the sequence we will use the EBI website, and for analysis we will use other
online tools.

Task 1: Extracting the DNA sequence We want to work on the Escherichia coli strain K12 recA
gene, which is in ENA entry with accession number V00328.

Task 1: instructions Go to the European Nucleotide Archive: [Link] and


search for the accession number “V00328”. Choose the sequence result with the correct
accession number. First view the text entry to get more details about the gene.

1. What is the UniProt accession number for the corresponding protein? Back on the entry page,
save the sequence in FASTA format. Do not save it in MSWord, save it as a text or .fsa file. Now
go to [Link] which provides a set of freely available online tools for sequence
analysis. Select “Composition” and then the “Genomics %G~C Content Calculator”. Calculate
the nucleotide composition of your sequence.

2. What are the numbers of each of the bases in your sequence and GC content? Is it GC or AT
rich?

Assessment Criteria

Task/ An excellent Very good Good answer Poor answer Very poor
Questio answer answer answer
n
Task 01 05 04 03 02 01
Q1 Excellently Very good Good attempt to Poor attempt to A very poor
introduced one of attempt to introduce one of introduce one of attempt
database and one of introduce one of database and one database and
analysis tool database and one of analysis tool one of analysis
of analysis tool tool
Q2 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
introduced one of attempt to introduce one of introduce one of attempt
database and one of introduce one of database and one database and
analysis tool database and one of analysis tool one of analysis
of analysis tool tool
Task 02 05 04 03 02 01
Q1 Excellently Very good Good attempt to Poor attempt to A very poor
identified number attempt to identify number identifynumber attempt
of hits associated identify number of hits associated of hits
with colon cancer in of hits associated with colon associated with
human genome with colon cancer cancer in human colon cancer in
in human genome genome human genome
Q2 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
identified number attempt to identify number identify number attempt
of loci using identify number of loci using of loci using
Genebank of loci using Genebank Genebank
Genebank
Q3 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
identified locus ID attempt to identify locus ID identify locus attempt
and position of identify locus ID and position of ID and position
MLH1 and position of MLH1 of MLH1
MLH1
Q4 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
identified the %ID attempt to identifiy the identifiy the attempt
of nucleotide identify the %ID %ID of %ID of
sequence for its of nucleotide nucleotide nucleotide
possible orthologs sequence for its sequence for its sequence for its
in mouse possible orthologs possible possible
in mouse orthologs in orthologs in
mouse mouse
Q5 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
identified the total attempt to identify the total identify the attempt
number of identify the total number of total number of
mutations of MLH1 number of mutations of mutations of
reported in human mutations of MLH1 reported MLH1 reported
gene mutation MLH1 reported in in human gene in human gene
database human gene mutation mutation
mutation database database database
Q6 2.5 2 1.5 1 0.5
Excellent attempt to Very good Good attempt to Poor attempt to A very poor
give the DNA attempt to give give the DNA give the DNA attempt
sequence of MLH1 the DNA sequence of sequence of
sequence of MLH1 MLH1
MLH1
Q7 2.5 2 1.5 1 0.5
Excellent attempt to Very good Good attempt to Poor attempt to A very poor
give the DNA attempt to give give the DNA give the DNA attempt
sequence of E. coli the DNA sequence of E. sequence of E.
mismatch repair sequence of E. coli mismatch coli mismatch
gene mutS. coli mismatch repair gene mutS. repair
repair gene mutS. gene mutS
Task 03 05 04 03 02 01
Q1 Excellently given Very good Good attempt to Poor attempt to A very poor
the accession attempt to give give the give the attempt
number, entry name the accession accession accession
and release date of number, entry number, entry number, entry
last modification. name and release name and release name and
date of last date of last release date of
modification. modification. last
modification.
Q2 05 04 03 02 01
Excellently given Very good Good attempt to Poor attempt to A very poor
the number of attempt to give give the number give the number attempt
amino acids, the number of of amino acids, of amino acids,
molecular weight amino acids, molecular weight molecular
and theoretical pI molecular weight and theoretical weight and
and theoretical pI pI theoretical pI

Q3 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
Calculated the total attempt to Calculate the Calculate the attempt
number of Calculate the total total number of total number of
negatively charged number of negatively negatively
residues and negatively charged residues charged
positively charged charged residues and positively residues and
residues and positively charged residues positively
charged residues charged
residues
Q4 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
Calculated the attempt to Calculate the Calculate the attempt
hydrophobicity of Calculate the hydrophobicity hydrophobicity
the MLH1 hydrophobicity of of the MLH1. of the MLH1
the MLH1
Q5 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
identified the attempt to identify the identify the attempt
number of identify the number of number of
peptides ,after number of peptides ,after peptides ,after
cleavage and the list peptides ,after cleavage and the cleavage and
of peptides with a cleavage and the list of peptides the list of
mass bigger than list of peptides with a mass peptides with a
1000 dalton with a mass bigger than 1000 mass bigger
bigger than 1000 dalton than 1000
dalton dalton
Task 04 05 04 03 02 01
Q1 Excellently Very good Good attempt to Poor attempt to A very poor
compared MLH1 attempt to compare MLH1 compare MLH1 attempt
and mutS sequence. compare MLH1 and mutS sequen and mutS seque
and mutS sequenc ce nce
e
Q2 05 04 03 02 01
Excellently Very good Good attempt to Poor attempt to A very poor
translated the above attempt to translate the translate the attempt
two gene sequences translate the above two gene above two gene
to protein sequences above two gene sequences to sequences to
sequences to protein protein
protein sequences sequences sequences
Q3 05 04 03 02 01
Excellently given Very good Good attempt to Poor attempt to A very poor
10 highest hits. attempt to give 10 give 10 highest give 10 highest attempt
highest hits. hits. hits

Q4 05 04 03 02 01
Excellently given Very good Good attempt to Poor attempt to A very poor
the pairwise attempt to give give the pairwise give the attempt
alignment and % of the pairwise alignment and % pairwise
sequence similarity alignment and % of sequence alignment and
of sequence similarity % of sequence
similarity similarity
Q5 10-7.5 7.5-5 5-2.5 2.5-1 01-00
Excellently given Very good Good attempt to Poor attempt to Inappropriate
the position of the attempt to give give the position give the answers.
CD, name of CD the position of the of the CD, name position of the
and Pfam ID CD, name of CD of CD and Pfam CD, name of
number. and Pfam ID ID number. CD and Pfam
number. ID number.

Q6 05 04 03 02 01
Excellently showed Very good Good attempt to Poor attempt to A very poor
the multiple attempt to show show the show the attempt
alignment of MLH1 the multiple multiple multiple
conserve domain alignment of alignment of alignment of
with 5 sequences MLH1 conserve MLH1 conserve MLH1 conserve
from the top of the domain with 5 domain with 5 domain with 5
CD alignment sequences from sequences from sequences from
the top of the CD the top of the CD the top of the
alignment alignment CD alignment

Marks obtained by the student


Marks obtained by
Total marks
Task Question Number the student for the
Allocated
answer provided
01 01 05
02 05
02 01 05
02 05
03 05
04 05
05 05
06 2.5
07 2.5
03 01 05
02 05
03 05
04 05
05 05
04 01 05
02 05
03 05
04 05
05 10
06 05
Total 100

Submission Guidelines
 Submission format: Report

 Paper Size: A4
 Words: 1000- 1500 words
 Printing Margins: LHS; RHS: 1 Inch
 Binding Margin: ½ Inch
 Header and Footer: 1 Inch
 Basic Font Size: 12
 Line Spacing: 1.5
 Font Style: Times New Roman
 Referencing should be done strictly using Harvard system

You might also like