0% found this document useful (0 votes)
18 views19 pages

RNA: Structure, Types, and Functions

Protomics

Uploaded by

na831032
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views19 pages

RNA: Structure, Types, and Functions

Protomics

Uploaded by

na831032
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

UNIT FOUR

Transcriptomics
RNA: Properties, Structure, Composition,
Types, Functions
RNA (Ribonucleic acid) is a single-stranded nucleic acid molecule and made up of
ribonucleotides.
 A ribose nucleotide in the chain of RNA consists of a ribose sugar, phosphate group, and
a base.
 In each ribose sugar, one of the four bases is added: Adenine (A), Guanine (G), Cytosine
(C), and Uracil (U).
 The base is attached to a ribose sugar with the help of a phosphodiester bond. As RNA
comprises many ribose nucleotides, the length of the chains of nucleotides can vary
according to their types or their functions.
 RNA thus differs from DNA, on the type of sugar used to make the molecule and
replacement of base Thymine in DNA with Uracil in RNA. Additionally, DNA is a
double-stranded molecule whereas RNA is a single-stranded molecule.
 The RNAs carrying the code for protein synthesis are called “coding RNAs” or
“messenger RNAs (mRNAs)”. Surprisingly, recent evidence revealed that very little of
our human genome sequences (less than 2%) could actually end up producing proteins.
However, most of the rest genome sequences are actively transcribed to generate the so-
called “non-coding RNAs (ncRNAs)”.
 These ncRNAs do not undergo translation to synthesise proteins but may hold the key to
broadening our understanding of gene regulation and human diseases. Many of them are
reported to serve as various regulatory elements in the genome, whereas most are still of
unknown importance to gene regulation.
What If Humans Suddenly Went Extinct?

Properties of RNA
 RNA is a single-stranded molecule and not a double helix, one of the consequences of
this, is that RNA can form a variety of three-dimensional molecular complexes than
DNA.
 RNA has ribose sugar in its nucleotides, rather than deoxyribose. These two sugars differ
from each other in the presence or absence of an Oxygen atom.
 Ribose sugar is also a cyclical structure consisting of 5 Carbon and one Oxygen just like
DNA. But the major difference is the presence of extra OH group in 2’ Carbon of
ribose which is absent in deoxyribose sugar.
 The OH group in 2’ Carbon makes the RNA molecule prone to hydrolysis.
 Some studies have also concluded that this chemical liability of RNA due to extra OH-
the group has led to DNA being the genetic storehouse as it lacks OH group in
2’Carbon making it more stable to hold information.
 RNA nucleotides carry the nitrogenous bases, Purines, and Pyrimidines. But in RNA in
place of Pyrimidine Thymine, Uracil is present which too forms hydrogen bonding with
Adenine just as Thymine does in DNA molecule.
Structure of RNA
Structure of RNA.
 RNA is a typical single-stranded biopolymer of ribonucleotides bonded with each other
via a phosphodiester bond.
 An RNA strand is synthesized in the 5’ to 3’ direction from a locally single-stranded
region of DNA.
 It has ribose sugars that are attached to four bases: Adenine, Guanine, Uracil, and
Cytosine. Ribose sugar has an extra OH- group in 2’ Carbon as compared to
deoxyribose sugar in DNA.
 This extra OH- group in RNA, has led them to be synthesized for short-term functions.
 The three-dimensional structure of RNA is critical to its stability and function.
 RNA being a single-stranded molecule can form a complex structure by allowing its
ribose sugars and bases to be modified on the action of cellular enzymes (that attach
chemical groups), to perform different functions.
 They are even capable of folding on themselves and showing intramolecular hydrogen
bonding between complementary strands, making it a double-stranded molecule to
exhibit specific function.
Composition of RNA
 RNA is a biopolymer of nucleotides bonded with each other via a phosphodiester bond.
 The nucleotide that makes up the RNA are also referred to as Ribose nucleotide due to
the presence of ribose sugar in their structure. Overall, RNA is composed of a ribose
sugar, phosphate, and nitrogenous base.
 Ribose sugar is a cyclical structure made up of five carbons and one oxygen atom. This
sugar contains two OH-groups at 2’ Carbon and 3’ Carbon.
 This ribose sugar is attached to a nitrogenous base via hydrogen bonding.
 There are four nitrogenous bases namely: Adenine (A), Guanine (G), Uracil (U), and
Cytosine (C).
 These nitrogenous bases pair complementarily with each other: G with C and A with U.
Types of RNA
Of many types of RNA, the three well known and most commonly discussed and found in almost
all organisms. These three types of RNA are:
1. mRNA (messenger RNA)
2. rRNA (ribosomal RNA)
3. tRNA (transfer RNA)
1. mRNA (messenger RNA)
mRNA (messenger RNA). Created with [Link]
 Messenger RNA (mRNA) carries the genetic code from DNA in a form that can be
recognized to make proteins. The coding sequence of the mRNA determines the amino
acid sequence in the protein produced. Once transcribed from DNA, eukaryotic mRNA
briefly exists in a form called “precursor mRNA (pre-mRNA)” before it is fully
processed into mature mRNA.
 This processing step, which is called “RNA splicing”, removes the introns—non-coding
sections of the pre-mRNA. There are approximately 23,000 mRNAs encoded in the
human genome.

 It is a single-stranded RNA molecule that is complementary to one of the strands of
DNA.
 mRNA is the version of the genetic materials that leave the nucleus and move to the
cytoplasm where responsible proteins are synthesized.
 This RNA has utmost importance during protein synthesis, when the ribosome moves
along this mRNA, it reads the base sequences and uses the genetic code to translate
them into specific proteins.
 These codes are in the form of triplet sequences of nitrogenous bases and are often
referred to as codons.
 As we now know that mRNA is responsible for transferring the genetic information into
ribosomes where by reading the base sequences on mRNA, the translation of proteins is
made possible, thus the name resembles its functions i.e., messenger RNA.
 We can even say mRNA is the molecule that uses genetic code present on a portion of
DNA and make proteins. If mRNA wouldn’t have existed then the information on
DNA could never be used by our body.
2. rRNA (ribosomal RNA)
 rRNA (ribosomal RNA).
 Ribosomal RNA is the catalytic component of ribosomes. In the cytoplasm, rRNAs and
protein components combine to form a nucleoprotein complex called the ribosome which
binds mRNA and synthesizes proteins (also called translation).
 It is a single-stranded RNA molecule found in cells that forms the part of the protein-
synthesizing organelle, Ribosome.
 It is synthesized inside the nucleus particularly in the nucleolus where rRNA coding
genes are present. The synthesized rRNA can be of varying sizes, commonly
distinguished as small and large.
 These newly synthesized rRNAs combine with ribosomal proteins and form smaller
subunits and larger subunits of ribosomes respectively.
 These rRNAs are vital in recognizing conserved regions of incoming mRNAs and tRNA
thus facilitating their binding and carrying out protein synthesis.
 Additionally, rRNA also has enzymatic activity (peptidyl transferase) and catalyzes the
formation of the peptide bond in between two aligned proteins/amino acids during
protein synthesis.
3. tRNA (transfer RNA)
tRNA (transfer RNA).
 Transfer RNA is a small RNA chain of about 80 nucleotides. During translation, tRNA
transfers specific amino acids that correspond to the mRNA sequence into the growing
polypeptide chain at the ribosome.
 It is a type of RNA molecule that helps to decode information present in mRNA
sequences into specific proteins.
 It is encoded by DNA in the cell nucleus and transcribed with the help of RNA
polymerase ΙΙΙ.
 The structure of tRNA folds upon itself and creates an intra complementary base pairing
which gives raise to hydrogen-bonded stems and associated loops that contains
nucleotides with modified bases.
 The structure in two-dimensional resembles a cloverleaf having three loops and an open
end. are usually 75-90 ribonucleotides in length.
 Each of these loops consisting of arms has a distinct name and function. The three-loop
consisting arms are namely: DHU or D arm, which has recognition site for specific
enzyme amino-acyl tRNA synthetase; T arm that consists of ribosome recognition site
and Anticodon arm that recognizes and bind to mRNA present in the ribosome.
 The open end with no loop is the site for attachment of amino acid, via 3’ OH bonding
with COOH- group of the amino acid.
 In general, tRNA reads the code on the mRNA sequence in Ribosome and translates
specific amino acid, it does so along the length of the mRNA and gives out a
polypeptide chain of amino acids (proteins) in association with other important
enzymes like aminoacyl tRNA synthetase and peptidyl transferase.
Some other types of RNA
Beyond the primary role of RNA in protein synthesis, several varieties of RNA exist that are
involved in post-transcriptional modification, DNA replication, and gene regulation. Some forms
of RNA are only found in particular forms of life, such as in eukaryotes or bacteria.

 Regulatory RNAs

A number of types of RNA are involved in regulation of gene expression, including micro RNA
(miRNA), small interfering RNA (siRNA) and antisense RNA (aRNA).

miRNA (21-22 nt) is found in eukaryotes, and acts through RNA interference (RNAi). miRNA
can break down mRNA that it is complementary to, with the aid of enzymes. This can block the
mRNA from being translated, or accelerate its degradation.

siRNA (20-25 nt) are often produced by breakdown of viral RNA, though there are also
endogenous sources of siRNAs. They act similarly to miRNA. An mRNA may contain
regulatory elements itself, such as riboswitches, in the 5' untranslated region or 3' untranslated
region; these cis-regulatory elements regulate the activity of that mRNA.

 Transfer-messenger RNA (tmRNA)

Found in many bacteria and plastids. tmRNA tag the proteins encoded by mRNAs that lack stop
codons for degradation, and prevents the ribosome from stalling due to the missing stop codon.

 Ribozymes (RNA enzymes)


RNAs are now known to adopt complex tertiary structures and act as biological catalysts. Such
RNA enzymes are known as ribozymes, and they exhibit many of the features of a classical
enzyme, such as an active site, a binding site for a substrate and a binding site for a cofactor,
such as a metal ion.

One of the first ribozymes to be discovered was RNase P, a ribonuclease that is involved in
generating tRNA molecules from larger, precursor RNAs. RNase P is composed of both RNA
and protein; however, the RNA moiety alone is the catalyst.

 Double-stranded RNA (dsRNA)

This type of RNA has two strands bound together, as with double-stranded DNA. dsRNA forms
the genetic material of some viruses.

 Small nuclear RNAs (snRNA; 150 nt):

Small nuclear RNAs are always associated with a group of specific proteins to form the
complexes referred to as “small nuclear ribonucleoproteins (snRNP)” in the nucleus. Their
primary function is to process the precursor mRNA (pre-mRNA).

 Small nucleolar RNAs (snoRNA; 60-300 nt):

Small nucleolar RNAs are components of small nucleolar ribonucleoproteins (snoRNPs),


which are complexes that are responsible for sequence-specific nucleotide modification.

 Piwi-interacting RNAs (piRNA; 24-30 nt):

Piwi-interacting RNAs bind the PIWI subfamily proteins that are involved in maintaining
genome stability in germline cells. Piwi-interacting RNAs also play a role in gametogenesis.

 MicroRNAs (miRNA; 21-22 nt):

MicroRNAs are small ncRNAs of ~22 nucleotides (nt) and the most widely studied class of
ncRNAs. These RNA species mediate post-transcriptional gene silencing through RNA
interference (RNAi), where an effector complex of miRNA and enzymes can target
complementary mRNA by blocking the mRNA from being translated or accelerating its
degradation. In humans, miRNAs are estimated to regulate the translation of >60% of
protein-coding genes.

 Long noncoding RNAs (lncRNA):

Long noncoding RNAs are a heterogeneous group of non-coding transcripts larger than
200 nt in size and make up the largest portion of the mammalian non-coding
transcriptome. It is estimated that more than 8,000 lncRNAs encoded in the human
genome. lncRNAs are essential in many physiological processes. To date, various
mechanisms of gene regulation by some lncRNAs have been reported, whereas most are
still of unknown function.
Functions of RNA
 The prime function of RNA is in protein synthesis.
 Without RNA, the information encoded in DNA could have never been transcribed to
make essential proteins that a cell needs to maintain its integrity.
 mRNAs have now been widely used in pharmaceutical industries to synthesize potential
vaccines.
 Moreover, mRNAs are now used to develop new categories of medicines
 mRNAs have made the formation of the cDNA library possible.
 rRNAs are structural units of Ribosomes, which are essential organelles during protein
synthesis.
 Ribozymes can help suppress the expression of specific mRNA by cleaving them out
without relying on the host’s machinery.
 Artificial antisense RNAs are capable of arresting protein synthesis by binding with the
mRNAs, which have contributed to the human’s ability to combat diseases and
mutations.
The transcriptome
The term transcript refers to the RNA strand produced when a gene is transcribed. RNA is a
more general term encompassing all types of RNA molecules, while a transcript specifically
refers to a complementary RNA copy of a DNA sequence generated through transcription. The
transcriptome encompasses all transcripts present in a biological system, and within the
transcriptome, there are multiple RNA species. Some of the most abundant RNA species include
mRNA, rRNA, tRNA, miRNA and lncRNA, each possessing distinct size, structure, or function.
Protein-coding transcripts (mRNA) are typically a few thousand nucleotides long and make up
most of the diversity of transcriptomes. Despite the high diversity of mRNAs, their numbers are
dwarfed by ribosomal RNA (rRNA), which can make up roughly 90%of the total numbers of
transcripts in cells . For mRNA quantification, it is often desirable to isolate the mRNA
molecules from the total transcriptome. Fortunately, a distinct structural difference between
mRNAs and other RNA species facilitates this separation. When mRNA molecules are
synthesized in cells, they typically undergo a series of post-transcriptional modifications that
involve shortening and extension. Polyadenylation is a process by which the RNA molecules are
extended with a chain of adenine bases known as the poly-A tail. In mammalian cells, the poly-
A tail is roughly 250 nucleotides long and it plays a role in the transportation, translation and
stability of the of the RNA molecule . Poly-A tails are mostly found on mRNA molecules and
some lncRNA molecules, and they can be selectively targeted using poly-A capture strategies to
enrich the protein-coding transcriptome.
1. From RNA to protein
The principal role of the protein-coding transcriptome is to convey genetic information from
DNA for protein synthesis. Proteins are the effectors that determine cellular structure and
function, and measuring proteins can therefore offer more informative insights into cells than
RNA measurements. Many cellular processes are governed by alterations in protein expression
and structure, and interactions between proteins. While transcriptome is sometimes used as a
proxy for the proteome to study processes that are governed by proteins, it is important to note
that the correlation between the two is often weak. There are several factors that contribute to
this weak correlation. First, mRNA translation can be enhanced or suppressed, disrupting the link
between transcription and translation. Second, the half-lives of proteins span from seconds to
days, whereas the half-lives of mRNA molecules are generally much shorter. This difference in
molecular turnover exacerbates the disconnect between their expression levels. However, the
transcriptome is a critical component of the flow of genetic information within cells, providing
valuable information about cellular characteristics and functions. Moreover, transcriptomic
profiling technologies currently offer high coverage, sensitivity, and scalability, making them
some of the most accessible methods for gene expression profiling. Advanced transcriptomic
methods provide exceptional resolution and can be utilized to profile single cells. In contrast,
obtaining proteomic data from individual cells has been a significant technical challenge. The
dynamic range of protein expression levels, which can span over seven orders of magnitude, is
one of the major obstacles to achieving this goal. mRNA on the other hand, has a more limited
dynamic range of only three to four orders of magnitude. Additionally, mammalian cells have a
more extensive range and abundance of proteins than mRNA molecules. A single mammalian
cell can contain billions of proteins, but less than half a million mRNA molecules. Furthermore,
proteins can be modified through post-translational modifications such as phosphorylation and
glycosylation, which increase the protein repertoire. Despite these obstacles, recent progress in
the proteomics field
suggests that new proteomics technologies are catching up to the level of sensitivity and
resolution currently achieved with transcriptomics technologies.
2. Transcriptomics technologies
Transcriptomics technologies have revolutionized genomics research by offering comprehensive
gene expression measurements at the RNA level. A range of biotechnological innovations form
the foundation of these technologies. This chapter aims to provide an overview of the
transcriptomics field and the existing experimental techniques.
2.1 A brief introduction to transcriptomics
Transcriptomics is the research field that investigates the transcriptomes of biological samples,
including cells and tissues. Many transcriptomics technologies rely on high-throughput
sequencing, which allows for the large-scale analysis of nucleic acid sequences.

One of the earliest transcriptomics methods, bulk RNA sequencing (bulk RNA-seq), measures
the combined transcriptomic profile of a large population of cells in a tissue sample. For two
decades, bulk RNA-seq was used extensively in comparative studies and large atlas projects to
classify and characterize biological samples. These atlas projects have generated considerable
amounts of data, providing a valuable resource for biological and clinical investigations.
Although bulk RNA-seq is a valuable technology in the field, it cannot profile the transcriptomes
of individual cells, which is crucial for understanding cellular characteristics and functions.

In 2009, a new variant of RNA-seq technology emerged which allowed researchers to profile
gene expression in individual cells. This technology would later become known as single-cell
RNA sequencing (scRNA-seq). Since 2009, scRNA-seq technologies have undergone rapid
development, with the introduction of more scalable alternatives that can profile hundreds of
thousands of cells. ScRNA-seq has quickly become an invaluable tool for biological research by
opening up new avenues for understanding cellular heterogeneity and gene expression
regulation. As a result of its impact in science, it earned the title of Method of the Year in 2013.
Despite its success, a major drawback of scRNA-seq is its inability to preserve information about
the spatial arrangement of cells within tissues, which is a critical element in many biological
processes.
A few years after the emergence of scRNA-seq, spatially resolved transcriptomics (SRT)
technologies were developed to profile gene expression in situ. These technologies measure
gene expression profiles with spatial resolution, placing biological phenomena within their
tissue context. Like scRNA-seq, SRT technologies have become indispensable tools in biological
research.
In recent years, they have demonstrated their impact across various fields, including
neuroscience, developmental biology, cancer biology and immunology. In 2020, SRT
technologies were recognized with the Method of the Year award, in recognition of their
substantial impact on biological research.

Figure 1– Trends in the field of transcriptomics. The popularity of transcriptomics methods is


reflected by the number of publications mentioning transcriptomics, single-cell transcriptomics
and spatially resolved transcriptomics over the last three decades. SRT is following a similar
trajectory as its predecessors, and is expected to gain in popularity over the next decade. The
search was conducted using a data analytics tool from Dimensions.
2.2 Bulk RNA sequencing
Transcriptomics technologies typically employ similar experimental approaches for capturing
and quantifying RNA. These principles were first established for measuring the transcriptomes
of larger cell populations using bulk RNA-seq. The experimental procedure for bulk RNA-seq
involves several steps, starting with the extraction of RNA from a biological sample. The
extracted RNA is then reverse transcribed into complementary DNA (cDNA) using a reverse
transcriptase (RT) enzyme. Primers are used to initiate this conversion, and a common technique
is to use primers that anneal to the poly-A tail of mRNA molecules to target transcripts
originating from protein-coding genes. Once the cDNA is synthesized, it is converted into
double-stranded DNA (dsDNA) and amplified using polymerase chain reaction (PCR). The
resulting amplified library consists of dsDNA fragments that can be decoded using high-
throughput sequencing technologies, such as Illumina sequencing.
Figure 2 – General principles of RNA sequencing. RNA is extracted from a tissue sample in
vitro. The mRNA is converted into complementary DNA (cDNA), then into double-stranded
DNA (dsDNA), followed by amplification by PCR. The amplified cDNA library can then be
decoded with high-throughput sequencing. In bulk RNA-seq, the quantified gene expression
levels represent the aggregated gene expression across all cells in the tissue sample.

Bulk RNA-seq provides measurements of expression levels aggregated across a larger


population of cells, which can limit the scope of biological inquiries that can be made. However,
bulk RNA-seq has been widely used in studies across large patient cohorts to identify molecular
processes associated with health and disease, leading to the identification of biomarkers and gene
expression patterns that have become widely utilized for disease classification, diagnosis, and
prognosis. While these endeavors have been significant, it is important to recognize that
biological samples are heterogeneous and composed of diverse cell types, each of which can
contribute to biological processes of interest. Therefore, to fully understand the mechanisms that
underlie biological processes, it is often more informative to examine the transcriptomes of
individual cells.

2.3 Single-cell transcriptomics


ScRNA-seq technologies utilize various strategies to quantify gene expression profiles in single
cells. The initial steps involve tissue dissociation to release the cells from their extracellular
matrix, followed by cell isolation. To isolate the cells, some of the most prominent techniques
utilize microfluidics, droplet-based systems, or picowell strategies. Once the cells are separated,
their RNA is reverse transcribed into cDNA, similar to the process used in bulk RNA-seq, but
with the
addition of cell-specific barcodes attached to the cDNA. The cDNA is subsequently amplified
and purified into a library that can be decoded with high-throughput sequencing. After decoding,
the cell-specific barcodes can be used to trace back transcripts to individual cells.

Initially, scRNA-seq technologies could only handle a few cells at the time and with limited
sensitivity. However, owing to numerous innovative solutions, modern technologies such as 10X
Genomics Chromium, Drop-seq and inDrop can profile hundreds of thousands of cells in a single
experiment. The scalability of modern scRNA-seq technologies has become essential for
extensive atlas efforts that have produced a wealth of publicly available scRNA-seq data. These
atlases provide detailed classification and characterization of cell types across different organs,
tissues and diseases.

ScRNA-seq has widespread applications, with key uses including the identification and
characterization of cell types and cell states, reconstruction of differentiation trajectories,
modeling of gene-expression dynamics, cell-cell communication, and comparative analyses
across biological conditions.
2.4 Spatially resolved transcriptomics
Although scRNA-seq has greatly advanced our understanding of cellular function, the spatial
organization of cells remains a relatively unexplored area. Studying cells outside of their tissue
context makes it difficult to identify the mechanisms that underlie the observed differences
between them. An important question that often arises in scRNA-seq analysis is whether distinct
transcriptomic profiles are intrinsic characteristics of cells acquired during an earlier stage of
their lifespan, or if they arise in response to changes in their environment. SRT technologies
overcome many limitations with single-cell RNA sequencing by profiling gene expression levels
directly in tissue sections. Various experimental methods have been proposed to achieve this
goal, but they differ significantly in fundamental aspects, such as spatial resolution and
multiplexing capacity. SRT technologies can be broadly classified into imaging-based and in situ
capture-based methods.

Imaging-based techniques represent the category with the highest spatial resolution as they can
resolve the precise position of individual RNA molecules. A major weakness with these methods
is the reduced multiplexing capacity, which isoften limited to a few hundred targets due to
optical limitations in the imaging process. In situ capture-based SRT technologies targets the
entire transcriptome, but their designs impose restrictions on the spatial resolution. There is a
clear tradeoff between the spatial resolution and the multiplexing capacity between imaging-
based and in situ capture-based SRT methods, and the selected method dictates the physical scale
and transcriptomic complexity at which biological processes can be studied.
Figure 3 – General principles of single-cell and spatially resolved transcriptomics.
B. Initially, a tissue sample undergoes dissociation to isolate individual cells. The cells are then
separated, for example, into lipid droplets or well plates. By using barcoding schemes, the
origin of transcripts can be traced to individual cells. B) In spatially resolved
transcriptomics, a common approach is to place tissue sections onto specially designed
capture arrays that track the position of transcripts. Each measured expression profile is
accompanied with a coordinate, enabling the mapping of gene expression levels to specific
locations within the tissue section.
It is worth mentioning that apart from spatial resolution and multiplexing capacity, there are
other aspects of SRT technologies which influence their applicability and popularity. These
include the cost of equipment and reagents, the training required to operate them, turn-around
time per sample, user-friendly analysis tools, options for multimodal analysis and their
commercial availability.
Figure 4 shows the number of published SRT datasets according to the museum of
SpatialTranscriptomics (2023), clearly indicating that the Visium technology currently dominates
the field [. The broad range of available SRT technologies necessitates a more detailed
discussion of their respective limitations and advantages. The following sections explore some of
the most popular SRT methods, as well as their fundamental differences in design and chemistry.

Figure 4 – Overview of SRT datasets. A) SRT technologies are ordered by the number of
published datasets. B) Organs studied with SRT technologies, ordered by the number of
published datasets. The colors represent the species that the datasets originated from. Data were
obtained from the Museum of Spatial Transcriptomics.
2.4.1 In situ capture techniques
In situ capture technologies, also called sequencing-based technologies, utilize barcoding
strategies to tag transcripts with spatial barcodes that contain positional information. These
methods utilize microarrays composed of densely packed units, sometimes referred to as spots,
beads, or pixels. For simplicity, we will refer to these units as spots. The diameters of spots vary
depending on the technology, ranging from 220 nm to 100 μm. Given that the average diameter
of a mammalian cell is approximately 10 μm, measurements are taken in the span of sub-cellular
to supracellular spatial resolution. Each spot is covered with DNA oligonucleotides consisting of
a unique spatial barcode and a short oligo-dT probe. Once a tissue section is positioned on the
array, the extracellular matrix and cellular membranes are enzymatically permeabilized, leading
to the release of RNA that spontaneously diffuse down to the array surface. Polyadenylated RNA
molecules hybridize to the oligo-dT probes on the spots, whereafter complementary cDNA
strands can be synthesized via reverse transcription. This process concatenates the cDNA of the
captured transcript to the spatial barcode, thereby combining the spatial and genetic information
into a single cDNA molecule. These cDNA molecules are then amplified and purified to produce
a library that can be decoded with high throughput sequencing. Owing to the poly-A capture
strategy, in situ capture-based technologies are unbiased and do not require selecting target gene
sets.

Figure 5– Resolution and multiplexing capacity in SRT technologies. A) Comparison of spot


sizes for 8 in situ capture SRT technologies. Stereo-seq achieves sub-cellular spatial resolution
with its nanoball-patterned array. DBiT-seq and slide-seq have a spot size that is comparable to
the size of an average mammalian cell. Visium captures 1-10 cells on average, depending on cell
density and tissue type. The multiplexing capacity of all in situ capture methods is transcriptome
wide. B) Comparison of multiplexing capacity for 18 imaging based SRT technologies. The
multiplexing capacity ranges between 100 to 10,000 target transcripts. Imaging-based techniques
can measure the position of individual transcripts. The data was obtained from the Museum of
Spatial Transcriptomics.

There are several in situ capture technologies available that achieve single-cell or sub-cellular
resolution. Slide-seq V2 and its predecessor Slide-seq uses randomly barcoded spots, each with a
size of 10 μm, that can measure the transcriptomic profiles of single cells or a mixed
transcriptomic profile from a few cells. Stereo-seq uses nanoball-patterned arrays where the
spots have an average diameter of 220 nm, and a center-to-center distance of 500 nm or 715 nm.
This design results in a density of 400 spots per 100 μm 2, making it the method with the highest
spatial resolution to date . DBiT-seq utilizes two parallel microfluidic channels to deliver DNA
probes to tag RNA molecules and proteins in situ, enabling joint analysis of RNA and proteins at
a spatial resolution of 100 μm.
Other examples of in situ capture technologies with sub-cellular spatial resolution are Seq-scope
and PIXEL-seq. These high-resolution SRT methods can measure RNA at the sub-cellular or
cellular level, but the RNA that reaches the surface and gets converted into cDNA only
represents a small fraction of the total tissue transcriptome. This leads to low capture efficiency,
which poses a challenge for data analysis as the data becomes sparse. Another limitation,
although not well-studied, is the effect of lateral diffusion of transcripts. For instance, in Stereo-
seq, the estimated lateral diffusion was reported to be 6.84 μm, significantly higher than the
center-to-center distance of adjacent spots (500-715nm), thus reducing the reliability of the
positional information [68]. To address the challenges associated with low sensitivity and high
lateral diffusion, a common strategy is to aggregate measurements from multiple neighboring
spots. Although this approach
lowers the spatial resolution, it allows for data exploration and analysis at different spatial scales.

The most widely used in situ capture technology is the Visium gene expression profiling
platform. The method also deserves additional attention as it represents the foundation for the
research presented in the scientific articles included in this thesis. The earliest version of the
technology was developed by Stahl et al. in 2016, and used a microarray with approximately
1,000 densely packed circular spots, each with a diameter of 100 μm. 10x Genomics later
commercialized and
further developed the technology into the Visium gene expression profiling platform, which
employs a similar microarray technology and chemistry as the earlier versions but with a smaller
spot size and a higher spot density. Each Visium spot measures 55 μm in diameter, providing
supra-cellular spatial resolution of approximately 1-10 cells per spot. The center-to-center
distance between adjacent spots is 100 μm and the capture area comes in two versions. A smaller
version can be used to analyze an area of 6.5x6.5mm2 with approximately 5,000 spots, while the
larger version covers an area of 11x11mm2 and approximately 14,000 spots. Unlike most other
in situ capture methods, the SRT dataset produced by Visium is complemented with a
histological image of the analyzed tissue section. This histological image offers valuable
information on tissue morphology and provides histological context to the transcriptomic
measurements.

Figure 6– The Visium SRT technology. A thin tissue section is cut from a tissue sample and
placed on a slide covered with densely packed spots. The tissue section is stained with
hematoxylin and eosin and imaged with a bright field microscope to obtain a histological image.
The tissue section is then permeabilized to release the mRNA which diffuses to the spots on the
surface of the slide and binds to DNA oligos. The DNA oligos in each spot contain a unique
spatial barcode that can be decoded to identify the position of each transcript. The output data is
a gene expression matrix with transcript counts for each detected gene, grouped by spots. The
transcriptomic profile of each spot represents a mixture of about 1-10 cells.
2.4.2 Imaging-based techniques
Imaging-based SRT technologies use microscopy to measure the position and abundance of
transcripts in tissue sections. These technologies share underlying principles with fluorescence in
situ hybridization (FISH) techniques, where selected transcripts are targeted using
complementary fluorescent probes. The fluorescent probes bind to their target RNA molecules
within the tissue section, and when the specimen is illuminated with light of specific
wavelengths, the fluorophores absorb the light and emit light of a longer wavelength. This
emitted light is measured with a microscope to record the precise location of the target
transcripts. Imaging-based SRT technologies such as MERFISH and seqFISH incorporate
various techniques to perform multiple rounds of hybridization and imaging, targeting different
transcripts in each round. Coupled with barcoding schemes, these technologies can target
hundreds to thousands of transcripts in a single tissue section. Another approach that relies on
microscopy imaging is In Situ Sequencing (ISS), which uses padlock probes to target transcripts.
Hybridized padlock probes are ligated and transformed into circular DNA molecules that are
amplified in situ via rolling circle amplification (RCA), which helps increasing the signal. The
amplified DNA molecules are decoded using sequential hybridization with fluorescent probes
and microscopy imaging.
2.4.3 ROI-based techniques
ROI-based techniques can be both imaging-based or sequencing-based and are designed to let
the user define a Region of Interest (ROI) to measure. The Nanostring GeoMX digital spatial
profiler is a widely used ROI-based technology, which adopts an innovative strategy to target
and capture RNA from tissue sections. The transcripts are first targeted with complementary
DNA oligos, which are connected to indexing oligo sequences by UV photocleavable linkers.
Regions of interest (ROIs) are selected based on visual staining patterns, using morphological or
molecular markers, where a single ROI can have a diameter of approximately 10-600 μm. UV
light is then directed onto the selected ROIs, which catalyzes the cleavage of the UV
photocleavable linkers, releasing the indexing oligos. The indexing oligos are extracted with
microcapillaries and measured with high-throughput sequencing or with the Nanostring nCounter
system. The method makes it possible for users to isolate and study areas of interest, for example
tertiary lymphoid structures, tumor compartments or stromal regions. However, defining
representative ROIs can be challenging and often requires guidance from experienced
pathologists.

You might also like