0% found this document useful (0 votes)
261 views9 pages

NCBI Database Overview and Usage Guide

Uploaded by

mairabilal78
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
261 views9 pages

NCBI Database Overview and Usage Guide

Uploaded by

mairabilal78
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Practical No.

01: Overview of NCBI Database


Introduction:
The National Center for Biotechnology Information (NCBI) is part of the United States
National Library of Medicine (NLM), a branch of the National Institutes of Health (NIH). It is
approved and funded by the government of the United States. The NCBI is located in Bethesda,
Maryland, and was founded in 1988 through legislation sponsored by US Congressman Claude
Pepper.
The NCBI houses a series of databases relevant to biotechnology and biomedicine and is an
important resource for bioinformatics tools and services. Major databases include GenBank for
DNA sequences and PubMed, a bibliographic database for biomedical literature. Other databases
include the NCBI Epigenomics database. All these databases are available online through
the Entrez search engine. NCBI was directed by David Lipman, one of the original authors of
the BLAST sequence alignment program and a widely respected figure in bioinformatics.
Website Url: [Link]
Username: itxnaveed942@[Link]
Password: Naveed94210.

Practical No.02: Searching & Retrieval of Data from NCBI.


Objective
To learn how to search and retrieve biological data using accession numbers from the National
Center for Biotechnology Information (NCBI) database.

Materials
- Computer with internet access
- Web browser
- NCBI account (optional)

Introduction
The NCBI provides a wealth of biological data, including DNA, RNA, and protein sequences.
Each entry in the NCBI database is assigned a unique accession number. This practical will
guide you through searching for and retrieving data using these accession numbers.

Procedure
Step 1: Accessing the NCBI Website

 Open your web browser.


 Go to the NCBI homepage at [Link]
Step 2: Searching Using an Accession Number

 On the NCBI homepage, locate the search bar at the top of the page.
 Enter the accession number (for example, `AH002765.2` for a Nucleotide sequence) into
the search bar.
 Click the search icon or press Enter.
Step 3: Understanding the Search Results

 The search results page will display information about the accession number entered.
 Look for the relevant entry, which will usually be the first result. This entry will contain
links to various databases like GenBank, RefSeq, and others.
Step 4: Accessing Detailed Information

 Click on the entry link to access the detailed view.


 In the detailed view, you will find multiple tabs, such as "GenBank," "FASTA,"
"Graphics," and "Analyze this sequence."
 Click on FASTA: This shows the sequence in FASTA format
Step 5: Downloading Data

 To download the sequence data, go to the "Send to" button located at the top right of the
detailed view page.
 Select the desired format (e.g., FASTA, GenBank) from the dropdown menu.
 Click on the "Create File" button to download the data to your computer.

Step 6: Using NCBI Tools for Further Analysis

 Explore the "Analyze this sequence" and click on Run BLAST for further analysis.
 BLAST (Basic Local Alignment Search Tool) allows you to compare your sequence with
other sequences in the database to find regions of similarity.

Practical No.03: Searching & Retrieval of Data from UNIPROT.


Objective
To learn how to search and retrieve protein data from the UniProt database using accession
numbers.
Materials Required
- Computer with internet access
- Web browser

Introduction
UniProt (Universal Protein Resource) is a comprehensive, high-quality, and freely accessible
database of protein sequence and functional information. Each protein entry in UniProt is
assigned a unique accession number which can be used to retrieve detailed information about the
protein.

Procedure
Step 1: Access the UniProt Website
1. Open your preferred web browser.
2. Navigate to the UniProt website: [Link]

Step 2: Understand the UniProt Interface


Familiarize yourself with the main sections of the UniProt homepage.
 Search Bar: For querying the database using various parameters including accession
numbers.
 Tools and Services: Additional tools provided by UniProt for data analysis.
 Help: Documentation and tutorials to assist users.

Step 3: Retrieve Data Using Accession Number

 Write the accession number `P05067` in the search bar and enter.
 The search results will display the protein entry associated with the accession number.

Step 4: Downloading Data


1. Click on the `Download` button located at the top right corner of the entry page and select the
desired format.
2. Download data in FASTA (canonical) format.
Step 5: Using Additional Tools and Resources
1. Then run BLAST: For sequence alignment and similarity searches.

Practical No.04: Searching & Retrieval of Data from BLAST N


(Nucleotide)
Practical No. 02
1. After Run BLAST
2. Select BLAST N:
- On the BLAST homepage, click on the "Nucleotide BLAST" link under the "Basic BLAST"
section. This will take you to the BLASTN search page.
3. Locate the Search Box for Accession Number:
- On the BLASTN search page, locate the search box where you can enter the nucleotide
sequence, accession number, or other identifiers.
4. Enter the Accession Number:
- In the search box, enter the accession number of the nucleotide sequence you want to retrieve.
For example, enter "NM_001301717" (you can use any valid accession number for your search).

5. Set the Parameters (Optional):


- You can set various parameters for your BLAST search, such as the database to search against
(e.g., "nr" for non-redundant nucleotide sequences), algorithm parameters, and filters. Write the
organism’s name of the given accession number and enter.

 To check in the same organism  Not Exclude.


 To check in different organisms  Exclude

I didn’t select to Exclude the Homo sapiens. It means I’m checking in the same organisms.
6. Execute the BLAST Search:
- Scroll down to the bottom of the page and click the "BLAST" button to initiate the search.
The search might take a few moments to complete.

7. Review the Results:


Once the search is complete, you will be directed to the results page. The results page shows a
graphical summary, descriptions of the hits, alignments, and a link to the original nucleotide
sequence.

8. Retrieve the Nucleotide Sequence Data:


- In the "Descriptions" section, select the entry that is most highly similar to the accession
number you entered. Click on the accession number link to view detailed information about the
nucleotide sequence.

- This will take you to the NCBI nucleotide database page for that sequence. Here, you can see
the complete nucleotide sequence, as well as additional information such as the gene name,
organism, and related sequences. I selected the 1st highest similar entry

9. Results of the Retrieved Data:


Write the results by noting down the following details:
Query length  20535
Accession Number  AH002765.2
Gene Name  Homo sapiens retinoblastoma susceptibility protein gene
Organism  Homo sapiens
Bits Score  4340 (Exact Score 2350)
Number of Matches  23
E value  0.0
Identities  100%
Gaps  0%

Common questions

Powered by AI

Accurately setting parameters is crucial when conducting BLAST searches as it directly affects the sensitivity, specificity, and computational efficiency of the results. Parameters such as the choice of database, algorithm (e.g., BLASTN for nucleotides or BLASTP for proteins), and filter options determine the quality of the output alignments. Setting organism-specific search parameters can help target specific phylogenetic queries, whereas selecting appropriate word sizes and gap penalties optimizes match performance. Precise parameter settings improve relevance of hits, enabling more meaningful biological insights and reducing false positives, thereby enhancing research accuracy .

BLAST enhances the analysis of biological sequences retrieved from NCBI by allowing researchers to compare a given sequence against a large database of known sequences to identify regions of similarity. This alignment process can reveal homologous sequences, aiding in annotating gene functions, predicting protein structures, and studying evolutionary relationships. BLAST offers different algorithms tailored for specific needs, such as BLASTN for nucleotide sequences, optimizing searches based on parameters like database choice and organism inclusion or exclusion. Such comparative analysis accelerates hypothesis generation and experimental design in biological research .

NCBI and UniProt serve distinct yet complementary roles in biomedical research. NCBI, part of the National Institutes of Health, focuses on a wide array of biological data, including DNA, RNA, and protein sequences, facilitating research with tools like BLAST for sequence alignment. In contrast, UniProt specifically caters to protein sequence and functional information, providing comprehensive, high-quality data crucial for understanding protein functions and interactions. While NCBI's databases such as GenBank are pivotal for genetic and genomic research, UniProt excels in offering detailed protein data, essential for proteomics and studying protein structures and functions .

Using accession numbers benefits researchers by providing a unique and stable identifier for each entry within databases like NCBI and UniProt, ensuring consistent data retrieval over time. This unique identifier simplifies data sharing and referencing, facilitating efficient access to specific datasets for further analysis or validation. Accession numbers allow researchers to bypass extensive queries, enabling direct navigation to desired sequences or entries, thereby saving time and reducing errors associated with less direct search methods. Such organization aids in maintaining data integrity across various research projects .

Understanding the FASTA format is crucial when retrieving sequence data from databases like NCBI and UniProt because it is a widely used text-based format for representing nucleotide or peptide sequences. The simplicity and readability of FASTA format, which includes a header line followed by lines of sequence data, makes it ideal for computational analysis. This format is compatible with various bioinformatics tools, enabling researchers to conduct further analyses such as sequence alignment or similarity searches efficiently. Adeptness in using and interpreting FASTA format ensures accurate data handling and processing in bioinformatics workflows .

Freely accessible bioinformatics databases like NCBI and UniProt have far-reaching implications for global scientific research and collaboration. They democratize access to critical biological data, enabling researchers from diverse economic and geographical backgrounds to contribute to scientific discovery. This accessibility facilitates cross-border collaborations, encouraging the exchange of ideas and data that can accelerate medical and scientific breakthroughs. Moreover, open access supports educational initiatives by providing comprehensive resources for training new scientists, thereby broadening the scope and reach of scientific research globally .

Databases like NCBI ensure data accuracy and reliability through rigorous curation and validation processes, involving expert curation teams who review and update entries based on the latest scientific findings. These procedures are essential as they maintain the integrity and trustworthiness of the data used in biomedical research. Accurate data is crucial for ensuring reliable results in experiments and analyses, as any errors could lead to incorrect conclusions and potentially hinder scientific progress or lead to inefficient resource allocation. High data reliability supports sound scientific inquiry and informed decision-making in medical fields .

NCBI plays a pivotal role in advancing open scientific data sharing by providing free and public access to extensive biological and biomedical databases like GenBank and PubMed. These resources enable researchers worldwide to share and access critical data, catalyzing collaborative scientific discovery and innovation. By offering tools and online platforms for data analysis, such as BLAST, NCBI encourages transparency and reproducibility in research. Its commitment to open data sharing enhances global research efforts in genomics and bioinformatics, breaking down silos and fostering a collaborative scientific community .

The foundation of NCBI in 1988, sponsored by US Congressman Claude Pepper, significantly contributed to advancements in bioinformatics by establishing a centralized resource for biological and biomedical data. As part of the National Institutes of Health, NCBI developed crucial databases like GenBank for DNA sequences, which facilitated data sharing and advancements in sequence alignment tools such as BLAST. The availability of these resources allowed researchers worldwide to access and analyze vast amounts of biological data efficiently, fostering developments in computational biology and genomics .

Effective search and retrieval in the NCBI database using accession numbers involves multiple steps. Initially, one must access the NCBI website and enter the accession number in the search bar. The search results should be examined to find the relevant entry, often linking to databases like GenBank. It's crucial to understand the data within the detailed view, accessible via various tabs such as "GenBank" and "FASTA," where one can download data in desired formats. Finally, tools like BLAST can be used for further sequence analysis by comparing with other sequences to find similarities .

You might also like