Overview of DDBJ Database Functions
Overview of DDBJ Database Functions
The primary distinctions between DDBJ’s Nucleotide Sequence Submission System (NSSS) and Mass Submission System (MSS) lie in their use cases and submission capabilities. NSSS is designed for small-scale submissions, handling individual or small sets of nucleotide sequences. It allows detailed information such as contact details, hold dates, and references to be submitted alongside the sequences . In contrast, MSS is intended for large-scale or complex submissions that involve more than 1024 sequences, sequences with more than 30 features, or particularly long sequences (>500kb). It necessitates separate sequence and annotation files .
DDBJ plays a pivotal role in advancing bioinformatics by providing a comprehensive and freely accessible nucleotide sequence database that supports scientific research worldwide. By issuing unique accession numbers for sequence data and maintaining synchronized communications with sister databases, DDBJ ensures a unified repository that is essential for comparative genomics and molecular studies . Additionally, DDBJ promotes bioinformatics education through training courses and software development, which equips researchers with the necessary tools for effective data analysis . Consequently, its contributions expand the capabilities of researchers and facilitate breakthroughs in understanding genetic diversity, genomic evolution, and biotechnology applications.
In the DDBJ Nucleotide Sequence Submission System (NSSS), templates are used to streamline the submission process by providing predefined formats for specific types of sequence data submissions. Users can select templates relevant to their samples, such as those for bacterial sequences, which standardizes the information required and ensures completeness and accuracy in data submission . Templates facilitate the submission of multiple sequences and ensure consistency across submissions, making the process more efficient for researchers and enhancing data integrity and reliability throughout the database .
The International Nucleotide Sequence Database Collaboration (INSDC), which includes the DNA Databank of Japan (DDBJ), National Centre for Biotechnology Information (NCBI), and European Bioinformatics Institute (EMBL), ensures data consistency by exchanging data submitted at any of the three databases. This means that the same data is available across all three, maintaining consistency across the databases .
DDBJ contributes to nucleotide sequence data accessibility and usability through a variety of tools and activities. It provides data retrieval tools such as getentry, which allows users to retrieve sequences using unique identifiers like accession numbers and gene names, and ARSA for a more comprehensive sequence and annotation search . Additionally, DDBJ develops software like WINA for data analysis and conducts bioinformatics training courses to enhance users' competence in using these tools . These activities collectively improve data accessibility and usability, making it easier for researchers to access and analyze sequence data.
Accession numbers play a crucial role in the management of sequence data at DDBJ by providing a unique identifier for each sequence entry. When researchers submit nucleotide sequences to DDBJ, an accession number is assigned, which ensures that each sequence can be uniquely identified and retrieved . This facilitates efficient data management, enables data traceability across international databases through INSDC partnerships, and supports data sharing and replication within scientific communities .
The DDBJ Sequence Read Archive (DRA) accommodates next-generation sequencing data by storing records of output generated by these machines, covering the primary analysis phase. As next-generation sequencing produces large volumes of data, the DRA provides a dedicated resource for submission and retrieval, utilizing tools like MetaDefine for web-based data submission . It is an important resource as it enables the storage and analysis of comprehensive sequencing data that is critical for modern genomic research, integrating with similar archives at NCBI and EBI through the INSDC .
The establishment and development of DDBJ's infrastructure have been supported through both domestic and international collaborations. Domestically, the DDBJ was established in 1986 within the National Institute of Genetics (NIG), Japan, backed by Japan’s Ministry of Education, Culture, Sports, Science and Technology. In 1995, the Center for Information Biology was founded within NIG to further support its operations . Internationally, DDBJ's functioning and growth are monitored by an international advisory committee with representatives from Japan, Europe, and the USA, ensuring alignment with global standards and collaborations within the INSDC for data consistency and sharing .
The DDBJ Read Annotation Pipeline offers cloud computing-based functionality that allows users to perform data analysis and annotation of raw sequencing data from next-generation sequencers submitted in the DRA. Its functionalities include Basic Analysis, which involves sequence mapping and de novo assembly, and High-level Analysis, which involves analytical processes for automatic and manual annotations like SNP detection. This pipeline enhances these processes by leveraging NIG's cloud computing resources, facilitating efficient and accessible sequence analysis .
Data exchange within the INSDC databases, including DDBJ, presents challenges such as ensuring consistency and accuracy across databases while managing the sheer volume of data submissions from around the world. However, this exchange is significant, as it allows researchers from any location to access comprehensive and unified data resources, aiding in global scientific progress . This challenge is met by daily data exchanges within the consortium, ensuring a synchronized data repository. For DDBJ, this means not only maintaining domestic operations efficiently but also collaborating closely with NCBI and EMBL to uphold the integrity of the shared repository .