0% found this document useful (0 votes)
22 views22 pages

DNA Digital Data Storage Overview

The document discusses DNA as a potential medium for digital data storage. It begins by explaining how DNA could theoretically store vast amounts of data in a very small physical space due to its ability to encode information in its nucleotide base pairs. It then provides an overview of the basic process of DNA digital storage, which involves encoding binary data as DNA nucleotide sequences and physically synthesizing DNA molecules to represent that data. The rest of the document outlines several chapters that would be included in a report on DNA digital storage, covering topics like how it works, past experiments conducted, advantages like long-term archival storage, challenges, and conclusions. It positions DNA storage as a potential future technology for very dense, durable long-term archiving of large amounts

Uploaded by

sf4432t
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
22 views22 pages

DNA Digital Data Storage Overview

The document discusses DNA as a potential medium for digital data storage. It begins by explaining how DNA could theoretically store vast amounts of data in a very small physical space due to its ability to encode information in its nucleotide base pairs. It then provides an overview of the basic process of DNA digital storage, which involves encoding binary data as DNA nucleotide sequences and physically synthesizing DNA molecules to represent that data. The rest of the document outlines several chapters that would be included in a report on DNA digital storage, covering topics like how it works, past experiments conducted, advantages like long-term archival storage, challenges, and conclusions. It positions DNA storage as a potential future technology for very dense, durable long-term archiving of large amounts

Uploaded by

sf4432t
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Report

1
Certificate

2
ABSTRACT

Digital data has changed the use and access of information.


Everyday lot of data is produced and this requires high-
density storage devices which can retain values for a long
time[1]. Deoxyribonucleic acid (DNA) can be potentially used
for these purposes as it is not much different from the
conventional method used in a computer. DNA can be used as
a robust and high-density storage device even under
unfavourable conditions[2]. Theoretically, one can encode 2
bits per nucleotide in DNA which can store 455 exabytes per
gram maximum data in single-stranded DNA (ssDNA)[3]. In this
paper, the method described can be used to store text data in
DNA by compressing, storing multiple copies along with
providing security to data.
Deoxyribonucleic Acid (DNA) is seen as a very important
medium for such purposes essentially, because it is similar to
the sequential code of zeroes and ones in a computer. This
field (DNA Computing) has evolved to become a topic of
interest for researchers since the past 10 years, with major
breakthroughs in its path. Seeming to come Straight out of
science fiction, “a coin-sized device could store the entire
information as the whole Internet”. The analyzed data from
the researches reveals that just 4 grams of DNA can Store all
the information that the world can produce in a year.

3
Content
Si no. TITLE. Page no:

1. CHAPTER 1: INTRODUCTION 5

2. CHAPTER 2: WHAT IS DNA DIGITAL STORAGE. 7

3. CHAPTER 3: LITERATURE REVIEW 8

4. CHAPTER 4: THE AGE LIMIT OF DNA STORAGE 9

5. CHAPTER 5: DNA DIGITAL STORAGE. 10

6. CHAPTER 6: WORKING OF DNA STORAGE 11

7. CHAPTER 7: EXPERIMENTS 12

8. CHAPTER 8: ADVANTAGES. 17

9. CHAPTER 9: CHALLENGES. 18

10. CHAPTER 10: CONCLUSION 19

11. CHAPTER 11: REFERENCES. 20

16. CHAPTER 12: REFERENCES. 22

4
Chapter 1
Introduction
DNA digital data storage is the idea of encoding binary data in
a DNA molecule and strand. It is a cutting-edge theory of data
storage that represents the new frontier of where technology
is going in the 21st century along with other major theoretical
advances like quantum computing.
The tremendous power of DNA digital data storage is largely
linked to the potential to fit tremendous amounts of data into
extremely small storage spaces. Scientists estimate that
practically infinite amounts of data can be stored in several
grams of DNA by translating the binary data into the four
categories of DNA proteins in the strand, and physically
creating DNA molecules to match. This process of physical DNA
construction is what DNA digital storage is based on, and is still
in a very theoretical stage. Although scientists have become
able to manipulate DNA and even build it, the idea of DNA
digital storage is still in its infancy and being evaluated
according to its theoretical use cases.
DNA is a very robust material and it has a long
shelf life. The Information stored in DNA can be recovered
even after Thousands of years. As long as the DNA is stored in
dry, dark And cold conditions, DNA can be stored for a long
time. By Using Polymerase Chain Reaction techniques, it is
possible to Get as many copies as required. Thus, copying of
data can be Done easily and many copies of data can be
obtained. As DNA Can retain information for centuries, DNA

5
can be used for Long-term storage. Due to high density, the
DNA can store a Large amount of data in very small space. As
in approach, the data is stored in long virtual DNA molecule
but encoding is done using synthetically prepared short DNA
strand. Short strands will allow to easily manipulating data. It
is possible to read simultaneously and randomly read files
stored in DNA. Also, compression technique is used to
compress data without any loss. The 4 nucleotides of DNA
used in the model are Adenine which will be denoted as A,
Cytosine as C, Guanine as G and Thymine as T.

6
Chapter 2
What is DNA Digital Data Storage
Data is stored in binary digits (1s and 0s) in traditional
computing. In DNA data storage, the four nucleotide bases (A,
C, G, T) store and encode data. Information is stored in
permutations of three nucleotides bases, called [Link]
digital data storage is the process of encoding and decoding
binary data to and from synthesized strands of [Link] DNA
as a storage medium has enormous potential because of its
high storage density, its practical use is currently severely
limited because of its high cost and very slow read and write
[Link] June 2019, scientists reported that all 16 GB of text
from the English Wikipedia had been encoded into synthetic
[Link] 2021, scientists reported that a custom DNA data
writer had been developed that was capable of writing data
into DNA at 18 Mbps.

7
Chapter 3
Literature review
The capacity of a medium to store information is usually
measured by the Shannon information. Since the DNA
molecule is a heterogeneous polymer composed of a linear
chain of deoxyribonucleotide monomers each adopting one of
four bases A, T, C and G, the specific arrangement (i.e.
sequence) provides a certain amount of information.
According to the definition of Shannon information, the
maximal amount of self-information (H) that a single base can
hold is
Where P(i) represents the probability of base I to occur at any
position, and log represents the base 2 logarithm as the bit
(binary unit) is usually used as a measurement of digital
information [21]. If and only if the four bases are equally likely
to occur, that is, Pi = ¼, each base pair in the DNA molecule can
provide the largest information capacity, i.e. 2 bits. The
dependence of self-information on base distributions is given
in Table 1, where a is the ‘probability distribution deviation’,
that is, the difference between the frequency at which the
base appears and the average frequency of 0.25.

8
Chapter 4
Reasearch History
In 1953, Watson and Crick published one of the most
fundamental articles in the history of biology in Nature,
revealing the structure of DNA molecules as the carrier of
genetic information [8]. Since then, it has been recognized that
the genetic information of an organism is stored in the linear
sequence of the four bases in DNA. In just a decade, many
researchers had proposed the concept of storing specific
information in DNA [9–11]. However, the concept failed to
materialize because the techniques for synthesizing and
sequencing DNA were still in their infancy.

In 1988, the artist Joe Davis made the first attempt to construct
real DNA storage. He converted the pixel information of the
image ‘Microvenus’ into a 0–1 sequence arranged in a 5 × 7
matrix, where 1 indicated a dark pixel and 0 indicated a bright
one. This information was then encoded into a 28-base-pair
(bp) long DNA molecule and inserted into Escherichia coli.
After retrieval by DNA sequencing, the original image was
successfully restored. In 1999, Clelland proposed using a
method based on ‘DNA micro-dots’ like steganography to store
information in DNA molecules. Two years later, Bancroft
proposed using DNA bases to directly encode English letters,
in a way similar to encoding amino acid sequences in DNA .

9
Chapter 5
The Age Limit Of DNA Storage
DNA molecules naturally decay with a characteristic half-life
[64,65], leading to a gradual loss of stored information. The
half-life of DNA highly correlates with temperature and the
fragment length. For example, Allentoft concluded that a DNA
molecule of 500 bp has a half-life of 30 years at 25°C, which
extends to 500 years for a fragment of 30 bp. Interestingly,
fossils provide empirical evidence of DNA’s stability over
thousands of years [65]. In this case, stability is significantly
improved by low temperatures and waterproof
environments]. Indeed, at −5°C, the half-life of the 30-bp
mitochondrial DNA ragment in bone is predicted to be 158 000
years. Other studies have explored packaging materials for
DNA molecules and have demonstrated impressive stability.
Grass et al. encapsulated solid-state DNA molecules in silica
and showed that they had better retention characteristics than
pure solid-state DNA and DNA in liquid environments. Judging
by first-order degradation kinetics, they concluded that it
could survive for 2000 years at 9.4°C or 2 million years at
−18°C, surpassing all potential quantitative data storage
materials invented to date. It is reasonable to expect a long
lifetime for data stored in DNA even at room temperature,
which makes DNA storage especially suited for cold data with
infrequent access. Further research may extend the lifetime of
DNA storage over the duration of human civilization with
minimal maintenance.

10
Chapter 6
DNA STORAGE SYSTEM
We imagine DNA digital data storage as the last level of a
Deep storing hierarchy, giving very dense and durable
Storage with access times of many hours to days. DNA
Synthesis and sequencing can be made arbitrarily parallel,
Making the necessary read and write bandwidths available.
We now detail our proposal of a system for DNA based
Storage with random access support.

11
Chapter 7
WORKING OF DNA DIGITAL DATA

A digital data in DNA should come across 5 levels and are as


Follows: Coding, Synthesis, Storage, Retrieval and Decoding.

❖ ENCODING

[Link] frequency table of characters of the data.


[Link] Huffman tree of non-repeating nucleotides for
Encoding is generated as follows:
▪ Each node in the tree will have 3 children.
▪ The weights of branches of children will depend on

12
The incoming weight of parent.
▪ On the off chance that the weight of incoming
Branch of a parent is A, at that point C represents to
The leftmost child, G represents to the middle child
And T represents the rightmost child.
▪ On the off chance that the weight of incoming
Branch of a parent is C, at that point G represents
The leftmost child, T represents the middle child and
A represents the rightmost child.
▪ On the off chance that the weight of incoming
Branch of a parent is G, at that point T represents
The leftmost child, A represents the middle child and
C represents the rightmost child.
▪ On the off chance that the weight of incoming
Branch of a parent is T, at that point A represents
The leftmost child, C represents the middle child and
G represents the rightmost child.
▪ T will be considered to be an incoming weight for Root.

[Link] split the whole data into overlapping segments of


100 nucleotides with an offset of 50 nucleotides from
Previous.

13
[Link] pairs of segments starting from the 1st segment.

5. Index each pair from 0 to 107 and after 107, start from 0
Again.
6. Reverse complement 2nd segment in each pair.
7. The index will be of 4 nucleotides long. The index is
Encoded by a combination of nucleotides in a sequence of A,
C, G, T such that no 2 consecutive nucleotides same.
Example: 0=ACAC, 1=ACAG, 2=ACAT.
8. Prepend A and append C to the 1st segment of the pair.
9. Prepend T and append G to the 2nd segment of the pair.
10. Each segment is now synthesized to actual DNA strand of
Length 106 nucleotides. If the length of the code of a
Character is 1, then to avoid repetition of nucleotides, 1 more
Nucleotide is added in the code of the character.

❖ DECODING

1. The decoding process is simply the reverse of the


Encoding process.
2. The 1st nucleotide of DNA will tell whether the DNA is the
1st or 2nd segment of the pair or whether the data is reverse

14
Complemented or not and directionality of strand.
3. If 1st nucleotide is A then:
▪ Remove 1st nucleotide.
▪ Next 4 nucleotides will tell us about segment Number.
▪ Next 100 nucleotides will be data.
▪ The last nucleotide can be used for confirmation of The
type of segment.
4. If 1st nucleotide is C then:
▪ Reverse whole segment.
▪ Remove 1st nucleotide.
▪ Next 4 nucleotides will tell us about segment Number.
▪ Next 100 nucleotides will be data.
▪ The last nucleotide can be used for confirmation of

The type of segment.


5. If 1st nucleotide is G then:
▪ Reverse whole segment.
▪ Remove 1st nucleotide.
▪ Next 4 nucleotides will tell us about segment Number.
▪ Reverse complements next 100 nucleotides.
▪ These 100 nucleotides will now be data.
▪ The last nucleotide can be used for confirmation of
The type of segment
6. If 1st nucleotide is T then:
▪ Remove 1st nucleotide.
▪ Next 4 nucleotides will tell us about segment Number.
▪ Reverse complements next 100 nucleotides.

15
▪ These 100 nucleotides will now be data.
▪ The last nucleotide can be used for confirmation of
The type of segment.
7. If TTTT sequence is found, this will denote the end of the
File. The new character will start from next nucleotide.
8. Now by using the same Huffman tree, data can convert
the

Data into original characters. It is possible to generate


Different Huffman tree for different files or single Huffman
Tree for whole data. This will compress the data and
Decoding cannot be done unless one has the original tree. As
Specific orientation nucleotides have been used in the
Strands, it is possible to read double number segments in the
Same number of indexes. The user can read the strand from
Any direction.

16
Chapter 8
EXPERIMENTS
To display the feasibility of DNA storage with random access
Abilities, we encoded four picture files using the two
Encodings. The files changed in assess from 5kB to 84kB. We
Joined these files and sequenced the ensuing DNA to
Recuperate the files. This section depicts our experience with
The synthesis and sequencing procedure, and presents comes
About exhibiting that DNA storage is functional and that
Random access works. We utilized the results of our
Experiments to illuminate the design of a simulator to Perform
more experiments exploring the design space of Data
encoding and durability.

17
Chapter 11
CONCLUSION

DNA-based storage has the potential to be the ultimate


Archival storage solution. It is extremely dense and durable.
Thus using DNA for data storage, it is possible to store huge
Amount of data in very less size. As DNA can hold data for
Many years, it is conceivable to store data for quite a while.
By utilizing this strategy, data is packed and the security to
The data is given. Parallel reading of files is also possible,
Enabling users to read multiple files at the same time. This
Technique maintains two copies of data. Hence in case of data
Damage, its copy can be used to read data. In the case of any
Errors, while encoding the data the error is restricted to that
Particular file and no other file is affected due to that error.
This technique can be used for all kind of files by making
Minor changes to adapt to the type of file. This technique can
Be used to store big data in very small space with little
Computational overhead. This method is scalable and can be
Used to store large files too. Also, multiple copies can be
Made easily. This method can be used to store information in
Archival systems or big data. Instead of using conventional

18
Storage devices which have less capacity to store data, DNA
Based storage method can be used in distant future to store
Data in a secured manner and for a long time storage and can
Solve the problem of limited space.

19
Chapter 12
REFERENCES
[1] J. Gantz, D. Reinsel. “Extracting value from chaos”,
International Data Corporation (IDC), Framingham (2011).
[2] C. Bancroft, T. Bowler, B. Bloom, C. T. Clelland. “Long
Term Storage of Information in DNA Science” (2001).
[3] George M. Church, Yuan Gao, Sriram Kosuri. “Next
Generation Digital Information Storage in DNA” (2012).
[4] M. Burrows and D. J. Wheeler, “A block-sorting lossless
Data compression algorithm,” Digital System Research
Center, USA, 1994.
[5][Link], [Link], [Link], and [Link], “Some
Possible codes for encrypting data in DNA”, 2003.
[6] D. A. Huffman, “A method for the construction of
Minimum-redundancy codes,” 1952.
[7] L. Adleman. Molecular computation of solutions to
Combinatorial problems. Science, 266(5187):1021–1024,
1994.

20
Chapter 9
Advantages
Although it has some challenges, DNA storage holds
immense potential for the future of storage. As research and
development continue, we may see DNA storage become a
viable option for storing and archiving large amounts of data
in the future. From preserving cultural heritage to advancing
scientific research, DNA storage could significantly impact how
we store and access information.
DNA storage offers several advantages over traditional
methods, from its potential for massive data storage capacity
to its durability and long-term stability. Let’s explore these
advantages in more detail
➢ Capacity and Durability:
DNA molecules can store vast amounts of data in a small
space, making it ideal for organisations that need to store large
amounts of data in a compact form.
➢ Longevity:
DNA storage has the potential to last for thousands of years if
stored under the right conditions, making it an ideal solution
for long-term data storage
➢ Energy Efficiency and Sustainability:
Traditional storage solutions are not as energy efficient as DNA
storage. Once the data is encoded into DNA and synthesised,
it can be stored at room temperature without additional
energy or cooling.

21
Chapter 10
The Challenges
While DNA storage technology holds great promise for the
future of data storage, we must overcome several challenges
before becoming a widely adopted solution. Here are some of
the main challenges of DNA storage:
➢ Cost and Accessibility:
DNA storage is much more expensive than traditional storage
methods like hard drives, flash drives, or cloud storage. The
cost of synthesising and storing DNA molecules, along with the
cost of reading and decoding the data, can be prohibitive for
many organisations. For instance, it has been estimated that it
would cost approximately $1 trillion to store one petabyte of
data (equivalent to one million gigabytes) using DNA synthesis
and sequencing technologies, according to a 2021 report from
the Massachusetts Institute of Technology (MIT).
➢ Speed of Writing and Reading Data:
In today’s fast-paced world, speed is everything, but one major
challenge is precisely the speed of writing and reading data to
and from [Link] methods of writing data to DNA are
slow and can take several hours or even days.
➢ Scalability:
the other challenges that researchers and scientists face is
scalability. The cost and time required for synthesising DNA are
still relatively high, so scaling up DNA production for large-
scale data storage applications can be challenging

22

You might also like