Repetitive DNA in Human Genome Review
Repetitive DNA in Human Genome Review
[Link] OPEN
Repetitive DNA sequences playing critical roles in driving evolution, inducing variation, and
regulating gene expression. In this review, we summarized the definition, arrangement, and
structural characteristics of repeats. Besides, we introduced diverse biological functions of
repeats and reviewed existing methods for automatic repeat detection, classification, and
1234567890():,;
masking. Finally, we analyzed the type, structure, and regulation of repeats in the human
genome and their role in the induction of complex diseases. We believe that this review will
facilitate a comprehensive understanding of repeats and provide guidance for repeat anno-
tation and in-depth exploration of its association with human diseases.
R
epetitive DNA sequences (repeats) are patterns of nucleic acids that occur in multiple
copies throughout the genome1. Both eukaryotic and prokaryotic organisms contain a
certain proportion of repeats in the genome2–4, particularly mammalians, in which repeats
account for 25–50% of their entire genome (Supplementary Fig. S1). For instance, about 50% of
the human genome consists of repeats5, while roughly 4% of human genes harbor transposable
elements in their protein-coding regions6. Because many of these repeats (~89.5%) are located
within introns, they have been erroneously assumed to be non-functional7. However, increasing
research indicates the significant impacts that repeats in coding and noncoding regions can have
on evolution, gene expression regulation, and variation induction8–10. For example, when repeats
are present in the coding region they get translated canonically. Not only can non-coding repeats
be translated by a non-canonical mechanism11, but even the telomeric repeat RNAs can get
translated12. Moreover, recent studies have shown that such repeats are closely related to a
variety of diseases, such as genetic disorders (e.g., Hemophilia), neurological diseases (e.g., poly-
Q diseases), and cancers (e.g., endometrial, stomach and colorectal cancers)13–15. A glossary
table (Supplementary Table S1) used to explain acronyms/terminologies in this study is shown in
Supplementary Note 1.
DNA sequences can be categorized into three groups according to their recurrence
frequency16, as shown in Fig. 1(a). The first group is composed of high-frequency repeats, also
known as satellite DNA sequences (satDNAs), which are found in various regions of the
chromosomes, including pericentromeric, subtelomeric, and interstitial regions. These sequences
typically form constitutive blocks of heterochromatin that are essential components of structures
such as centromeres and telomeres17. The length of satDNA repeating units can vary from a few
base pairs to over 1 kilobase pairs, forming arrays that can span up to 100 megabases and be
repeated over 106 times, making up ~8–10% of the human genome18.
The second group comprises moderate-frequency repeats that are typically 500–300,000 base
pairs in length and repeated between 10 and 105 times, accounting for ~30% of all repeats19.
These repeats are further classified into two subcategories: (A) microsatellites and minisatellites
(VNTR), and (B) dispersed repeats, which are primarily made up of transposable elements
1 Computational Bioscience Research Center (CBRC), Computer, Electrical and Mathematical Sciences and Engineering Division, King Abdullah University of
Science and Technology (KAUST), Thuwal 23955, Saudi Arabia. 2 Department of Endocrinology, Yichang Central People’s Hospital, The First College of
Clinical Medical Science, China Three Gorges University, 443000 Yichang, P.R. China. ✉email: [Link]@[Link]
(a)
DNA sequences
(b)
DNA Sequence
Tandem repeats
Interspersed repeats
(c) The general structure of DNA transposons (g) The distribuon of repeats in the human genome
~3kb-10kb
TSD TIR 5'-UTR Transposase 3'-UTR TIR TSD
Terminal Inverted Repeats
Target Site Duplicaon
~0.7-4kb
SVA
TSD (CCCTCT)n Alu-like VNTR SINE-R poly(A) TSD
3%
(TEs)20. It is worth noting that many moderate-frequency repeats control gene expression22,23. Approximately 40–50% of the total
have been implicated in gene expression regulation21. human DNA sequences are single-copy DNA sequences, meaning
The third group comprises unique, single-copy DNA sequen- that about half of the human genome is composed of unique and
ces, which do not share homology with any other sequences in non-repetitive sequences.
the genome. Examples of such sequences in the human genome According to the arrangement of repeating units, repeats can
include protein-coding genes (e.g., the globin, ovalbumin, and silk be classified into two types: tandem repeats (TRs) and inter-
fibroin genes), non-coding RNAs, and regulatory elements that spersed repeats24, as depicted in Fig. 1(b). Interspersed repeats,
Fig. 1 General classification of repeats, the typical structure of TEs and TRs, and the proportion of various types of repetitive elements in the human
genome. Sub-graph (a): Classification of repeats in the human genome. Sub-graph (b): Arrangement and characterization of repeats in the human genome.
Sub-graph (c): Typical structure of DNA transposons, in which TIR and TSD respectively represent the terminal inverted repeat and target site duplication.
Sub-graph (d): Typical structure of non-LTR retransposons, in which the color blocks represent the protein domains contained in each family, and the gray
block represents the non-coding regions. Sub-graph (e): Typical structure of retrovirus-like LTR retrotransposons, in which LTR represents the long terminal
repeat. Sub-graph (f): Typical structure and distribution of TRs in the human genome. Sub-graph (g): Proportion of TRs and active TEs in the human
genome. Specifically, LINE-1 and LINE-2 retransposons are represented by L1 and L2 respectively, while SINE-VNTR-Alu retrotransposon and Mammalian-
wide interspersed repeats are represented by SVA and MIR. The color arrows represent the repetitive unit (or motif) of each kind of TR, and the light black
structure represents the chromosome.
Table 1 Classes and length distribution of tandem repeats in are located at the telomeres, consisting of 300–8000 precise
the human genome. CCCTAA/TTAGGG motifs and covering a range of 2–50 kb on
the end of the chromosomes32. Subtelomeric repeats are located
in the boundary of 100–300 kb between the telomere and the
Class of TRs Length of TR unit Length of TR array
remaining part of the chromosome, consisting of satellite-like
Telomeres ~6 bp ~10–15 kb sequences33. Type, length, frequency, and distribution of TRs in
Tandem paralogous the human genome are summarized in Table 1 and Supplemen-
rDNA ~43 kb ~3–6 Mb tary Table S2.
Segmental ~1–400 kb ~1kb–5Mb
duplications
Microsatellites ~2–6 bp ~10–100bp
Transposons. Transposons are classified into RNA and DNA
Minisatellites ~10–100bp ~100bp–20kb transposons, depending on their mode of transposition. RNA
Satellites transposons use a cut-and-paste mechanism, where the transpo-
Alpha satellite ~171bp ~0.2–8Mb sase enzyme excises the transposon from its original location and
Beta satellite ~68 bp ~60–80kb inserts it elsewhere in the genome via an RNA intermediate. DNA
Gamma satellite ~48–220bp ~11–121kb transposons also use a cut-and-paste mechanism, but they move
Satellite I ~17–25bp ~2.5kb directly as DNA and are excised from their donor locus and
Satellite II ~23–200bp ~11–70kb reinserted elsewhere in a conservative mechanism. This diver-
Satellite III ~5bp ~3.6kb gence results in various dissimilarities in their transposition
Satellite IV ~35bp ~25–530kb mechanisms and evolutionary trajectories. Typical structures of
Macrosatellites ~100bp–5kb ~300kb retrotransposons, transposons, and tandem repeats are illustrated
Megasatellites ~1–5kb ~400kb
in Supplementary Fig. S2(a),(b) and (c), respectively.
A glossary table (Supplementary Table S1) included in supplementary, presenting detailed DNA transposons, also known as Class II transposons, can be
explanations for all acronyms and terminologies utilized in the manuscript. classified into four super families based on their constituent
structures and transposition patterns: miniature inverted-repeat
also known as transposons or TEs, consist of DNA and RNA TEs (MITEs), Cryptons, Mavericks (or Polintons), and Helitrons.
transposons25. Generally, TRs refer to a sequence array formed by MITEs are non-autonomous transposons primarily found in the
the repeated occurrence of basic repeating units connected head- non-coding regions of plant and animal genomes34, with the
to-tail26 (Supplementary Note 2). TRs, especially satellite DNA, ability to alter gene structures and functions. Cryptons are a
are clustered in specific chromosomal regions such as cen- unique class of DNA transposons that use Tyrosine Recombinase
tromeres, tetramers, and telomeres, which play an essential role in (YR) to cut and reattach recombining DNA molecules35, allowing
cellular processes, including chromosome segregation, genome them to incorporate YR sequences and drive animal evolution.
organization, and chromosome end protection27. For example, Mavericks are large DNA transposons commonly found in
centromeres contain long tandem arrays of alpha-satellite repeats eukaryotic genomes, with 6 bp target site duplication (TSD)
that extend over millions of base pairs and are organized in a sequences and genes homologous to viral proteins36. Helitrons are
hierarchical manner. The tandem arrays span between 100 and recently discovered eukaryotic transposons present in many plant
5000 bp on different chromosomes, ranging from 0.2 to 10 Mb. and animal species37, which propagate through a rolling circle
Some of these arrays include 17 bp binding motifs for the mechanism but don’t generate terminal repeats or TSDs. DNA
centromere-specific DNA binding protein, which have been used transposons are characterized by terminal inverted repeat
to create synthetic human chromosomes28. sequences (TIRs), which are complementary to each other at
the left and right ends of the transposon. These transposons, also
known as jumping genes, can move and integrate into diverse
Tandem repeats. Tandem Repeats in the human genome can be genomic regions. Figure 1 (c) illustrates the general structure of
divided into the following subcategories: microsatellites, minisa- DNA transposons in genomes. DNA transposons, which make up
tellites, centromeric satellites, and telomeric and subtelomeric about 5% of the human genome38, are considered DNA fossils
repeats (Fig. 1(f) and Table 1). The difference between micro- because no family of them currently remains active in most
satellites and minisatellites is represented in their length and mammals, including humans39,40.
frequency of occurrence. Microsatellites are DNA sequences of RNA transposons, also known as retrotransposons or Class I
<5 bp units repeated in tandem and are most frequent in the transposons, can be classified into five super families based on
human genome29. Minisatellites are tandem repetitions of more their structures and transposition patterns: Long terminal repeats
than 5 bp units, and their frequency in the human genome is (LTRs), Long interspersed nuclear elements (LINEs), Short
relatively rarer than that of the former30. In the human genome, interspersed nuclear elements (SINEs), Dictyostelium intermedi-
centromeric satellites can be classified into the alpha-satellite and ate repeat sequence (DIRS), and Penelope-like elements
Satellite II/III. Among them, Satellite II/III comprises of various (PLEs)41,42. LTR retrotransposons are related to retroviruses
variations on the ATTCC motif31. Telomeric repeats (satellites) and have LTRs at their 5′ and 3′ ends, which likely originated
from ancient retroviral infections43. LINEs contain an internal hemophilia B58, through mechanisms, such as loss of function
promoter that drives the expression of transposition machinery, mutation, modulation of splicing, and deletions at the site of
including reverse transcriptase and an endonuclease44. SINEs insertion. The general structures of non-LTR retrotransposons are
depend on LINEs for their transposition, with specificity presented in Fig. 1(d). The type, family, and length distribution of
determined by their 5′ tails. Most SINEs are derived from tRNA, repeats, as well as a brief introduction to their biological
7SL RNA, or 5s RNA and have an RNA-Pol III promoter45,46. functions, are shown in Supplementary Table S3.
DIRS retrotransposons, which have tyrosine recombinase, differ The general structure of retroviruses and LTR retrotransposons
from integrases or endonucleases commonly used by retro- are similar59. Several LTR retrotransposons have similar open
transposons for site-specific genomic integration47,48. PLEs share reading frames (ORFs) to those of retroviruses, consisting of the
an ancestor with telomerase reverse transcriptases (TERTs) and gag and pol (pro) genes and, in some cases, env and other
have unique features in retroelement phylogeny49. In the accessory genes. The main difference between retroviruses and
phylogeny of reverse transcriptases (RTs), PLEs do not belong LTR is the presence of a functional envelope (env) gene in
to the LTR or non-LTR retrotransposon groups but form a sister retroviruses, which is absent or nonfunctional in
clade with TERTs. TERTs are major components of the LTRretrotransposons60. The general structure of the retrovirus-
telomerase complex that maintain the linear chromosome ends LTR is illustrated in Fig. 1 (e). No retrotransposable LTR
in most eukaryotes50,51. retrotransposons have been identified in the human genome, and
The RNA transposons in the human genome can be classified no LTR retrotransposon insertions have been collected in the
into LTR and Non-LTR retrotransposons. Non-LTR retrotran- database of human mutations. However, many elements belong-
sposons lack LTRs, but contain genes for reverse transcriptases, ing to the young human endogenous retroviruses (HERV) family,
RNA-binding proteins, nucleases, and sometimes the Ribonu- such as HERV-K (K denotes a lysine-tRNA-specific primer
clease H domain52. LINE and SINE are two remaining active binding site to initiate reverse transcription), have an individual
super families contained in non-LTR retrotransposons of the ORF domain in their structure capable of translation and
human genome, consisting of LINE1 (L1), Alu, and SINE-VNTR- production of functional proteins61. Furthermore, HERVs and
Alu (SVA), three active families (Table 2). Many studies have mammalian apparent LTR retrotransposons (MaLRs) are rem-
suggested that L1 may contribute to human cancers by mutating nants of ancient retroviral infections found within the human
specific oncogenes or tumor suppressor genes in somatic cells53. genome. These genetic components are notable for their up-
For example, there is evidence that APC tumor suppressor gene regulation after innate immune activation and are primarily
failure is caused by the L1 insertions, which may be an important regulated in the context of immunity (Table 2). Retroelements
factor in the development of colorectal cancer54. In addition, Alu and isolated LTRs, as part of molecular evolution, may benefit the
elements are retrotransposons specifically present in primate host by promoting plasticity and gene expression regulation (i.e.,
genomes that can regulate gene function by providing canonical via promoters and cis-regulatory sequences)62. The expression of
polyadenylation signals and play a critical role in the primate HERV-K envelope transcripts is typically undetectable in normal
genomic diversity, causing complex diseases55. For instance, human breast tissues but is detectable in most breast cancer
many complex human diseases, such as meningococcal disease, tissues63. Therefore, this expression pattern can be used as a new
venous thromboembolism, obesity, and breast cancer, are related disease biomarker in clinical diagnosis. The general structure and
to the structural variants caused by Alu insertions56. Currently, distribution of tandem repeats, and the percentage of TE families
SVA is more active than high-copy pseudogenes (e.g., processed in the human genome are illustrated in Fig. 1(f) and (g),
ribosomal pseudogenes), and SVA insertions may alter gene respectively. The proportion of the most abundant repeats in the
expression and cause several human diseases57. For example, SVA genomes of Humans, Rice and Drosophila is presented in
regulates the expression of related genes whose insertions have Supplementary Fig. S1.
been identified as a significant contributor to diseases such as X- Sequence analysis techniques such as de novo assembly,
linked dystonia-parkinsonism, Neurofibromatosis type 1, and multiple sequence alignment (MSA), sequencing error correction,
One type of repetitive element that is unique to the human genome is known as the Human Endogenous Retrovirus (HERV). HERVs are remnants of ancient retroviral infections that occurred millions of
years ago and became integrated into the human genome. They comprise ~9% of the human genome and are considered to be a type of transposable element.
variations between humans and chimpanzees since their evolu- expression of genes. Changes in the regulation of critical genes
tionary divergence. For another example, a recent research can have profound effects on cellular processes, development, and
suggests that L1 insertions can cause genomic rearrangements, disease susceptibility.
including deletions, inversions, and duplications, leading to For instance, a study has revealed that over 120 independent
structural variations in the human genome101. The specific TE insertions are essential contributors to human diseases,
relationship between genome rearrangements caused by TEs and including hemophilia, Dent disease, neurofibromatosis and
complex diseases is discussed in Supplementary Note 6. cancers102. The germline transposition rate for the Alu element
in humans is about 1 in 21 births103, while the corresponding
Transposable elements can act as insertional mutagens in germline value for the L1 element is about 1 in 95 births104. Historically,
and somatic cells. Mobile elements, such as L1, Alu, SVA and TEs have generally been considered transcriptional silencing in
HERV-K, are in charge of novel germline insertions, which may somatic cells. However, evidence indicates that active TEs are also
lead to genetic illness (Table 3) (Supplementary Note 6.1 to present in the somatic cells of various organisms. As an
Note 6.7). The primary mechanisms by which TEs act as inser- illustration, the expression and transposition of the L1 element
tional mutagens in germline and somatic cells are described have been identified in several somatic contexts, such as early
subsequently: embryos and specific stem cells105. Furthermore, HERV-K
elements have been implicated in insertional mutagenesis. Recent
Disruption of coding sequences: When a TE inserts within a studies have identified HERV-K insertions with potential
coding region of a gene, it can disrupt the reading frame, intro- mutagenic effects on nearby genes, including cancer-related
duce premature stop codons, or cause other structural changes. genes106 (Supplementary Fig. S7(b)). Human cancers have also
This disruption can lead to the loss of gene function or the exhibited somatic activity, with tumors able to pick up hundreds
production of truncated and non-functional proteins. In germline of additional L1 insertions. For instance, recent research has
cells, such mutations can be inherited and contribute to genetic highlighted the impact of L1 insertions in diseases such as cancer,
variation in subsequent generations. neurological disorders, and genetic syndromes107.
Alteration of regulatory elements: TEs can insert near regulatory Transposable elements can drive key coding and non-coding RNAs.
elements, such as promoters, enhancers, or insulators, and disrupt According to mounting evidence, TE insertions may serve as the
their function. This can result in the misregulation or aberrant building blocks for forming protein-coding genes and non-coding
The relationships between TEs and diseases were summarized from refs. 55,58,78,221. Similarly, the associations between TRs and diseases were summarized from refs. 222–224.
RNAs that can carry out the crucial physiological functions of Enhancer hijacking: TEs can integrate near enhancer regions,
cells108. For example, Rag1 and Rag2 are spectacular examples of affecting the binding of transcription factors and changing the
deeply conserved TE-derived genes that activate V(D)J somatic regulation of nearby genes.
recombination in the immune system of vertebrates109. As
another example, based on a mixed lncRNA annotation from Promoter modulation: TEs can also insert near gene promoters,
RNA sequencing and GENCODE (a scientific project in influencing the recruitment of transcriptional machinery and
genome research and part of the ENCODE scale-up project), a impacting gene expression levels.
study estimated that 41% of lncRNA nucleotides are derived from In the human genome, L1 elements have the potential to
TEs, and the majority of lncRNAs (about 83%) contain at influence transcriptional networks. Recent research has demon-
least one TE fragment110. The primary mechanisms by which strated that L1 retrotransposition can introduce novel regulatory
TEs drive key coding and non-coding RNAs are described elements, alter gene expression patterns, and contribute to
subsequently: cellular diversity116. Furthermore, Alu elements can also impact
transcriptional networks. Recent studies have highlighted their
Retrotransposition: TEs, particularly retrotransposons, can role in shaping tissue-specific gene expression, alternative
undergo a process called retrotransposition where they are tran- splicing, and influencing the expression of neighboring genes
scribed into RNA and then reverse transcribed back into DNA, through enhancer or promoter activities117. The diverse mechan-
leading to their insertion into new genomic locations. If these isms through which TEs influence host gene-regulatory networks
retrotransposed elements land within or near functional genes, can be broadly categorized into five classes: (1) introduction of
they can act as alternative promoters, enhancers, or splice sites, transcription factor binding sites, promoters, and enhancers, (2)
giving rise to new coding and non-coding RNA transcripts. This modification of 3D chromatin architecture, (3) production of
process can generate novel RNA molecules with potentially regulatory non-coding RNAs, (4) usage of TE-derived coding
functional roles in cellular processes. sequences as new transcriptional effector proteins, and (5)
secondary effects of TE silencing mechanisms118.
Co-option of regulatory elements: TEs often contain regulatory
sequences such as promoters, enhancers, and insulators. These
sequences can be co-opted by the host genome to regulate the Biological functions of tandem repeats. TRs are common fea-
expression of nearby genes or to shape the expression patterns of tures of both prokaryote and eukaryote genomes. For example,
non-coding RNAs. By providing alternative regulatory elements, more than one million distinct TRs are contained in the human
TEs can impact gene expression networks and contribute to the genome, many of which are highly polymorphic in sequence
production of key coding and non-coding RNAs. composition and copy number. TRs can be found in intergenic
The presence of TEs that drive key coding and noncoding regions and in both the non-coding and coding regions of a
RNAs in the human genome may be associated with certain variety of genes119–121. Moreover, TRs occur near or between a
diseases (Table 3). For instance, HERVs affect human health and series of genes and can affect the structure and function of DNA,
cause disease by encoding proteins, acting as promoters/ RNA, and proteins through specific mechanisms and produce a
enhancers or lncRNAs, accounting for about 9% of the human series of molecular and cellular consequences122. As an illustra-
genome111. HERVs can also have a direct effect via their proteins tion, many TRs are involved in biological functions in a copy
in the development of cancers. For example, by inducing cell-cell number-dependent manner, and there is evidence that TRs may
fusion or epithelial-to-mesenchymal transition, HERV envelope regulate the expression of nearby genes by altering their copy
proteins play a critical role in tumorigenesis and development in number123. In general, TRs are highly mutable and can be located
melanoma, endometrial carcinoma, and breast cancer112. in exons, introns, or intergenic regions, providing opportunities
Furthermore, HERVs can generate lncRNAs that promote cancer for the modulation of gene expression, as well as the structure and
proliferation, motility, and invasion. For example, in the study113, function of RNAs and proteins124. Expanded TRs usually cause
researchers have found that several HERVs-derived lncRNAs, various disorders, including autism spectrum disorder (ASD) and
such as UCA1, SAMSON, and BANCR, are involved in the cancers (Table 3 and Supplementary Table S5). The illustrations
processes of proliferation, motility, and invasion in bladder in Fig. 4 and Supplementary Fig. S5 highlight how TR can directly
cancer and melanoma. The relationship between transcriptional or indirectly affect the genome.
activation of HERV retrotransposons and human cancer is
summarized in Supplementary Note 6.7.
Transposable elements can alter transcriptional networks and Accelerate the evoluon and adaptaon
conduce to cis-regulatory DNA elements. Cis-regulatory DNA
elements (CREs) are regions of non-coding DNA that regulate
the transcription of neighboring genes. In addition, CREs are Play an essenal role in the structural stability
vital components of genetic regulatory networks. Some TEs of genec materials during the whole cell cycle
TRs
have evolved into CREs, whose function is to mimic host pro-
moters, enabling them to recruit host-encoded factors driving
their selfish transcription114. For instance, due to innate and
Cause gene families and funconal redundancy
adaptive immune responses, the immune system can protect
organisms from pathogens and foreign substances. During
evolution, some TE families, including many endogenous ret-
roviruses (ERVs), have the capacity to influence and shape Cause a range of disorders, and regulate
transcriptional networks. They can function as signaling gene expression in healthy individuals
molecules that regulate DNA elements and the immune
system115. The primary mechanisms by which TEs alter tran-
scriptional networks and conduce to cis-regulatory DNA ele- Fig. 4 How TRs affect the genome. Similar to TEs, TRs can also affect the
ments are described subsequently: genome in specific ways.
Tandem repeats can accelerate evolution and adaptation. TRs are chromosomes. Telomeres, with their repeated sequences and
often referred to as satellite DNA, which can be further classified associated proteins, form a protective cap that allows complete
into microsatellites or short tandem repeats (STRs) (motif length: replication of chromosome ends and prevents the loss of genetic
1–4 bp), minisatellites (motif length: 5–64 bp), and macro- information. Telomeric TRs, in conjunction with telomerase
satellites (motif length: several kp), according to the size of the enzyme activity, ensure the integrity and stability of the genome
repeated motifs125. For example, slipped strand mispairing is a during successive cell divisions.
mutation process that occurs during DNA replication, which is For instance, centromeres are chromosomal domains respon-
one explanation for the origin and evolution of repetitive DNA sible for the faithful transmission of genetic material during cell
sequences126. TRs, especially STRs, are extremely unstable in division. They are characterized by highly repetitive DNA regions
terms of length, sequence composition, and copy number, with and bound kinetochore proteins, and they are required for the
mutation rates typically 10–100,000 times higher than in other attachment of microtubules to the chromosomes during
parts of the genome127. These unstable repeats are found in up to mitosis135. An array of tandem repeats known as alpha-satellites
20% of eukaryotic genes and promoters, where they confer phe- is one of the crucial components of centromeres, and it plays a
notypic or functional variability on the cell surface and extra- vital role in maintaining the stability of human chromosomes.
cellular proteins and have pathological consequences. The Variations in alpha-satellites can impact the function of the
primary mechanisms by which TRs accelerate evolution and centromere136. In addition, telomeres consist of repeat sequences
adaptation are described subsequently: and are bound by multiple telomeric interacting proteins. In
mammalian cells, telomere DNA is composed of double-stranded
Rapid genetic variation: TRs undergo rapid changes in copy tandem repeats of TTAGGG, with terminal 3′ G-rich single-
numbers and lengths, creating genetic diversity that can drive the stranded overhangs. Telomeres are protected by protein com-
emergence of new traits. plexes, such as shelterin, which includes TRF1, TRF2, POT1, and
other proteins that interact with telomeres indirectly137. This
Gene regulation: TRs located in regulatory regions can influence protection distinguishes natural chromosome ends from acci-
gene expression, allowing for adaptive changes to occur in dental DNA breaks and prevents unwanted repair machinery
response to environmental pressures. activity on telomeres.
In the human genome, TRs are also frequently found in genes Furthermore, the 5′ and 3′ UTRs of genes are transcribed but
that control body morphology128,129. For example, compared usually not translated into proteins. However, they contain
with synteny blocks, evolutionary breakpoint regions in the various regulatory elements involved in post-transcriptional gene
human genome contain more base pairs associated with TRs, regulation, such as mRNA stability, localization, and translation
with AAAT being the most frequent motif130. These TRs within efficiency138. STRs within UTRs can contribute to gene
evolutionary breakpoint regions have the potential to facilitate regulation in the following ways: (1) Modulation of mRNA
and accelerate gene expression evolution and generate sufficient stability: STRs in the UTRs can impact the stability of mRNA
variability to drive the rapid evolution and adaptation of molecules. Changes in STR length may affect the folding of
organisms131. Furthermore, recent studies have shown that STR UTRs, leading to altered interactions with RNA-binding proteins
variations in immune genes, such as HLA loci, can shape immune and subsequent degradation or stabilization of mRNA. (2)
responses and contribute to adaptation to diverse Regulation of translation efficiency: UTRs can also influence
environments132. In addition, TRs located in regulatory regions translation initiation and efficiency. STRs located in the 5′ UTRs
can facilitate evolutionary adaptations. Recent research has can affect ribosome binding and start codon recognition, leading
suggested that expansion or contraction of STRs within to changes in translation rates and protein production. STR
regulatory regions can modulate gene expression and contribute variations in UTRs have been associated with complex traits and
to phenotypic variation and adaptive responses133. diseases. For instance, a recent study identified UTR STR
expansions associated with the risk of neurodevelopmental
Tandem repeats can play a critical role in the structural stability of disorders139.
genetic materials during the cell cycle. Within or around certain In addition, TRs can be transcribed into RNA molecules
specialized chromosomal regions (e.g., centromeres, telomeres, through the process of transcription, which is carried out by RNA
and subtelomeres), TRs may play crucial roles in the structural polymerases140. When these TRs are transcribed into RNA, the
stability of genetic materials during the cell cycle134. The primary resulting RNA molecules can exhibit structural features and
mechanisms by which TRs play a critical role in the structural functional implications. The structure of TRs in terms of
stability of genetic materials during the cell cycle are described transcribed RNA are as follows: (1) Transcribed RNA molecules
subsequently: derived from TRs retain the repetitive nature of the underlying
DNA sequence. (2) TR RNA can fold into various secondary
Replication fork stabilization: TRs, consisting of repeated DNA structures due to intra-molecular base pairing within the
sequences adjacent to each other, can stabilize the replication repetitive sequence. (3) TR-derived RNA molecules can serve
forks during DNA replication. The repetitive nature of TRs diverse non-coding RNA functions. For example, some TR RNAs
provides a stable template for DNA polymerases to bind and act as scaffolds for the assembly of ribonucleoprotein complexes
initiate replication. This stability prevents replication forks from or regulate gene expression through interactions with RNA-
stalling or collapsing, ensuring accurate and complete DNA binding proteins or microRNAs. (4) TR-derived RNA can engage
replication. TRs act as essential structural elements that con- in regulatory mechanisms such as RNA interference, where
tribute to the stability of genomic regions during the cell cycle. complementary TR RNA pairs with target mRNA to modulate its
stability or translation. TR RNA molecules can also influence
Telomere maintenance: Telomeres, specialized TRs located at the cellular processes by sequestering RNA-binding proteins or acting
ends of chromosomes, play a crucial role in maintaining genomic as decoys for regulatory factors. (5) Expansions or contractions of
stability. Telomeres protect the ends of chromosomes from TRs in transcribed RNA have been linked to various genetic
degradation, fusion, and recognition as DNA breaks. During each diseases. Abnormal TR RNA structures and interactions can
round of DNA replication, the conventional DNA replication result in functional consequences, including the sequestration of
machinery has difficulty fully replicating the ends of linear RNA-binding proteins, disruption of cellular processes, or
induction of toxic effects. These factors contribute to the structure, leading to changes in the accessibility of regulatory
pathogenesis of diseases141,142. elements and the recruitment of transcriptional machinery. The
variability in TR length and sequence can impact the affinity of
Tandem repeats can result in redundancy of gene families and transcription factors for binding sites, resulting in differential
functions. A gene family is a collection of many related genes that gene expression levels.
typically perform comparable biological tasks. Individual mem-
bers of clustered gene families are often responsible for achieving Epigenetic regulation: TRs can act as susceptible targets for epi-
specific phenotypes or functions in the overall mission143. Tan- genetic modifications, such as DNA methylation and histone
dem gene duplication is thought to have significantly contributed modifications. The length and sequence composition of TRs can
to the evolution of large gene families, genetic and morphological influence the degree of epigenetic regulation. Methylation of TRs,
diversity, and speciation in eukaryotes144,145. The primary for example, can lead to the formation of repressive chromatin
mechanisms by which TRs result in redundancy of gene families and transcriptional silencing. These epigenetic modifications can
and functions are described subsequently: have a profound impact on gene expression patterns and con-
tribute to the regulation of various cellular processes.
Gene duplication: TRs can undergo replication slippage during
DNA replication, leading to the expansion of the repeat region Alternative splicing: TRs within exons or introns can affect
and subsequent gene duplication. This process can result in the alternative splicing, a process that generates multiple mRNA
creation of additional copies of genes within the same genomic isoforms from a single gene. Variation in TR length can influence
region. The duplicated genes are often subject to variations, such the splicing process by altering the stability of RNA secondary
as point mutations or insertions/deletions, that accumulate over structures or serving as binding sites for splicing factors. This can
time, leading to divergence in their sequences and functions. This result in the inclusion or exclusion of specific exons, leading to
duplication and subsequent diversification of gene copies can the production of different protein isoforms with distinct func-
result in redundancy within gene families, where multiple genes tions or regulatory properties.
have similar or overlapping functions.
Expansion: The expansion of TRs can also cause a range of dis-
Divergent evolution: Over time, duplicated genes arising from orders, known as trinucleotide repeat expansion disorders. When
TRs can undergo divergent evolution. Mutations and genetic the size of certain TRs exceeds a threshold, it can lead to genomic
changes accumulate in each gene copy, resulting in alterations to instability and pathological consequences. The expanded TRs can
their coding sequences and regulatory elements. These changes exhibit a tendency for further expansion and accumulation in
can lead to functional divergence, where duplicated genes acquire subsequent generations, resulting in a dynamic and progressive
different functions or have differential expression patterns. As a increase in repeat length. The expanded TRs can interfere with
result, redundant gene copies can contribute to the expansion and gene function, leading to impaired protein production, altered
diversity of gene families, providing evolutionary opportunities protein structure, or disrupted cellular processes. Trinucleotide
for gene innovation and adaptation to new environmental or repeat expansion disorders include conditions like Huntington’s
physiological contexts. disease, Fragile X syndrome, and several forms of spinocerebellar
For example, the genes responsible for coding ribosomal RNA ataxia, among others. These disorders often display a correlation
(rRNA) are present in the human genome as numerous tandemly between the size of the repetitive expansion and the severity of the
arrayed copies. These ribosomal DNA (rDNA) repeats facilitate disease phenotype.
the production of abundant amounts of rRNA to satisfy the cell’s For example, Lynch syndrome is an autosomal dominant
constant requirement for ribosome production146. In mammals, disorder that increases the risk of developing colorectal cancer,
rDNA repeats are present in two types of tandem arrays, termed endometrial adenocarcinoma, and tumors of the small intestine,
the 5S and 47S (or 45S) arrays. The 5S rDNA repeats are located in stomach, ureter, renal pelvis, ovary, brain, and prostate. Research
one large tandem repeat array on chromosome 1 in humans. The in study155 has demonstrated that most (90%) colorectal cancer
47S arrays are located on the short arms of five acrocentric due to Lynch syndrome have microsatellite instability. In
chromosomes in humans (chr. 13, 14, 15, 21, 22)147. Research addition, researchers in study156 have revealed that one
conducted by the Chinese Academy of Sciences investigated the neurodegenerative disease in which microsatellite instability
impact of TR-mediated expansions and variations within the contributes to a substantial number of cases is amyotrophic
mucin gene family. These TR expansions and variations contribute lateral sclerosis (ALS), a rapidly progressive and uniformly fatal
to the redundancy and functional diversification of mucins, which motor neuron disease. Recent research indicates that TR
play important roles in various cellular processes148. polymorphisms can also regulate gene expression in healthy
individuals133. Furthermore, TR instability can lead to reduced
Tandem repeats can regulate gene expression, and their expansion gene expression, increased disease incidence, and enhanced
can cause a range of disorders. TR instabilities, especially micro- tumor aggression (Supplementary Fig. S7(c) and (d)). The
satellite instability, contribute significantly to causing gene association between tandem repeat instabilities and cancer,
expression variation in humans149, and numerous disorders such autism, as well as neurological disorders, is discussed in
as cancer, ASD, Huntington’s disease, various ataxias, motor Supplementary Note 6.8 and Note 6.9.
neuron disease, frontotemporal dementia, and fragile X syn-
drome, are associated with the expansion of TRs, particularly
Repeat detection
STRs150–154 (Table 3). The primary mechanisms by which TRs
Numerous computational methods have been proposed for
regulate gene expression, and their expansion can cause a range of
identifying repeats in genomes, which can be divided into
disorders are described subsequently:
homology-based, structure-based, de novo methods, and hybrid
Transcriptional modulation: TRs located within gene regulatory frameworks, as shown in Table 4 and Supplementary Fig. S8.
regions, such as promoters and enhancers, can influence gene
expression by affecting the binding of transcription factors. The Homology-based identification methods. Homology-based
presence of TRs can alter the three-dimensional chromatin methods identify repeats by finding subsequences similar to
11
REVIEW ARTICLE
12
Table 4 (continued)
Method type Method name Description/Characteristic Advantages/Disadvantages References
Disadvantages: (1) This algorithm usually consumes
vast computing resources (CPU, memory, and disk
space) and has a long run time. (2) The detection
accuracy of the algorithm is usually unsatisfactory.
EDTAg The EDTA package is specifically designed to Advantages: (1) It demonstrates robustness across 205,231
minimize false discoveries in raw TE candidates, plant and animal species based on empirical evidence.
enabling the creation of a high-quality, non- (2) It is capable of deconvoluting nested TE
redundant TE library for comprehensive whole- insertions, which are commonly observed in highly
genome TE annotations. These annotations repetitive genomic regions.
REVIEW ARTICLE
‘Hybrid frameworks’ refer to detection tools that adopt multiple detection strategies, and they usually cannot be clearly distinguished into the above three typical types. ‘EDTA’ is the abbreviation of the extensive de novo TE annotator. ‘RepeatMod2’ is the abbreviation of
RepeatModeler2.
a[Link]
b[Link]
c[Link]
d[Link]
e[Link]
f[Link]
g[Link]
h[Link]
known repeats, which must rely on algorithms for comparing otherwise, terminate. RepeatFinder186, RepeatScout187, ReAS188,
similarity between sequences, such as the hidden markov model and Generic Repeat Finder (GRF)189 are representative of this
(HMM)-based comparison algorithm, and specific databases, class of approaches. The third class of methods includes
such as RepBase157, Dfam158, msRepDB159, REXdb160, and RepARK190, REPdenovo191, RepAHR192, and RepLong193, which
Pfam161. RepeatMasker ([Link] is a rely on de novo sequence assembly and community detection in
representation of such tools, which uses Dfam or RepBase as the sequence similarity network to identify repeats (Supplementary
backend library and RMBLAST ([Link] Table S8). Among these four tools, the first three obtain repeats
[Link]) as the aligner. RMBLAST and Dfam are a new by performing assembly of high-frequency reads or k-mers
aligner and database specially developed by RepeatMasker team (Supplementary Fig. S9(a),(b),(c),(d) and (e)). The last method
for repeat detection based on the existing aligner BLAST162 constructs the similarity network by getting the overlaps between
([Link] and database RepBase long reads, and then use the community discovery algorithm to
([Link] Both RMBLAST and Dfam get the repeats (Supplementary Fig. S9(f)). A detailed introduc-
have become gold standards in the field of repeat annotation. tion to the de novo identification methods is shown in Supple-
Typical homology-based detection methods also include mentary Note 7.1.3.
Censor163, TESeeker164, Greedier165, and T-lex166 (Supplemen-
tary Table S6). The advantages of homology-based methods lie in
Tandem repeat and their expansion identification methods.
their accuracy and the ability to discover families with a small
Several tools are available for detecting TRs and their expansions,
number of copies. Their disadvantage is that they cannot be used
such as mreps194, Tandem Repeats Finder (TRF)195, T-REKS196,
to discover new repetitive sequences that are not collected in
TRASH197, EnsembleTR198,199, RExPRT200, GangSTR201,
homology databases. A detailed introduction to homology-based
ExpansionHunter200, ExpansionHunter De novo202, Straglr203,
methods can be found in Supplementary Note 7.1.1.
and STRling204. Among them, mreps excels by detecting all types
of tandem repeats in an entire genomic sequence simultaneously.
Structure-based identification methods. Repeats, especially TEs, It incorporates a resolution parameter to identify fuzzy repeats
have specific structures, such as the structure of a protein, or non- with variations within the repeated units. TRF uses sequence
coding domains, and differ in the presence and size of the TSD, a alignment and statistics to detect consecutive repetitive motifs. It
short, direct repeat generated on both flanks of a TE upon gives detailed information about identified repeats, including
insertion167. Structure-based methods rely on prior knowledge of positions, consensus sequence, length, and alignment scores. This
structural features of known repeats collected in the library and information is valuable for genome analysis, gene mapping,
employ a heuristic algorithm to identify repeats in genomes. investigating structural variations, and understanding repetitive
Typical structure-based identification methods include elements in biology and evolution. T-REKS operates by dividing
LTRharvest168, MASiVE169, MGEScan-LTR170, TE-greedy- the input sequence into overlapping k-mer segments, where k is a
nester171, SINE-Finder172, SINE_scan173, AnnoSINE174, user-defined parameter. Then, it employs the k-means clustering
FINDMITE , MUST , detectMITE , MITE-Hunter34,
175 176 177 algorithm to group similar k-mers together, identifying potential
MITE-Digger178 and, MITE Tracker179 (Supplementary TRs. EnsembleTR and GangSTR, developed by the Gymrek Lab,
Table S7). The advantages of structure-based methods include are powerful tools in computational genomics and human
high detection efficiency and lower false-positive rate, and the genetics. EnsembleTR takes VCF files with TR genotypes for
detected repeats are easier to verify and classify. Their dis- multiple samples and generates a consensus set of genotypes.
advantages are that they cannot be used to identify repeats whose RExPRT is a machine learning tool used to differentiate patho-
structural features are unknown or whose structural features genic from benign TR expansions. GangSTR is a tool used for
cannot be obtained accurately and completely due to the insuf- profiling TRs across the genome using short reads. One notable
ficient precision and completeness of the input sequences. Thus, advantage of GangSTR is its ability to handle repeats that exceed
the detection integrity of such methods is often unsatisfactory. the read length. ExpansionHunter and ExpansionHunter De novo
Besides, structure-based methods are often designed for a parti- are two computational methods developed by Illumina Inc. to
cular class of transposons (e.g., LTRs, SINEs, and MITEs). locate both known and novel repeat expansions in short-read
Therefore their versatility is limited. A detailed introduction to sequencing data. Straglr is a specialized tool designed to identify
structure-based detection methods is shown in Supplementary and genotype TR expansions using whole genome long-read
Note 7.1.2. sequences. STRling is a method for detecting new short TR (STR)
expansions from short-read sequencing data, even when no
corresponding STR is present in the reference genome.
De novo identification methods. The de novo methods are more
flexible than the other two classes of methods because they do not
require prior knowledge about the structure or similarity to Hybrid frameworks. The classification of methods mentioned
known repeats180, which can also be classified into three cate- above is based on the core technology utilized in each method.
gories based on the core technology that each method depends However, there are certain detection tools like Extensive de novo
on. The first class of methods includes Repeat Pattern Toolkit181, TE Annotator (EDTA)205 and RepeatModeler2206, which employ
RECON182, PILER183, LTRdigest184, and LongRepMarker185, multiple existing detection algorithms or strategies to perform
identifying repeats through MSA. The strategy of high-frequency repeat annotation. These tools cannot be easily classified into the
k-mers and space seed extension is used in the second category of above-mentioned three categories due to their unique approach
methods to identify repeats. The sequences to be detected are that incorporates multiple existing methods for repeat annota-
converted into k-mers of a certain length, and k-mers whose tion. For example, EDTA incorporates various tools, such as
frequency exceeds a certain threshold are chosen as seeds. Then, RepeatModeler and RepeatMasker, which employ homology-
the locations of these seeds in the genome are recorded, and the based methods, as well as TransposonPSI. In addition, it incor-
repeats are obtained by performing sequence extensions at both porates structure-based methods like LTRharvest and
ends of the genome. During the extension process, the detection LTR_retriever. RepeatModeler2 is another hybrid framework,
algorithm always judges whether the extended arrangements are that utilizes the de novo methods RECON and RepeatScout,
consistent across multiple genome locations. If yes, continue; along with the Dfam database and the alignment search tool
RMBLAST, to identify and model repetitive elements in DNA Databases that support automated repeat classification and
sequences. Performance comparisons between different repeat masking. An accurate and comprehensive repeat database is
detection methods are shown in Supplementary Tables S9–S32 of essential for the automated classification and masking of repeats in
the Supplementary Note 7.2. genomes. Three well-known nucleic acid libraries, RepBase, Dfam,
msRepDB, and three famous protein libraries, RepeatsDB, REXdb,
and Pfam, have been proposed to support the automated classifi-
Automated classification and masking of repeats cation and masking of repeats. RepBase ([Link]
Classification and masking are two necessary steps after the repbase/) is a database of prototypic sequences representing repe-
detection stage in the workflow of repetitive DNA sequence titive DNA from different eukaryotic species, which currently
analysis. Precise classification and comprehensive masking of contains more than 38,000 sequences of different families. Dfam
repeats are essential for analyzing their critical roles in genomes. ([Link] database is an open
The output of the detection stage consists of raw repeat consensus collection of TEs and genome annotations, which currently houses
sequences without any information about the type, structure, and 285,542 TE models across 595 species and incorporated into the
function. The purpose of classification is to classify unknown new version of RepeatMasker. msRepDB ([Link]
repeats into their main taxonomic branches (e.g., LTR, LINEs, [Link]/pages/msRepDB/[Link]) is the most compre-
SINEs, DIRS, PLEs, MITEs, Cryptons, Helitrons, Mavericks, hensive multi-species repeat database, which currently contains
Satellites, low complexity sequences, etc.), and to distinguish their TEs of more than 84,000 species. RepeatsDB ([Link]
structures and functions. The purpose of repeat masking is to [Link]/) collects protein structures of annotated TRs, which
mask the repeats in the genome of a specific sequencing sample provides users with the possibility to access and download high-
with the well-classified elements collected in the repeat database quality datasets either interactively or programmatically through
using pairwise sequence alignment algorithms, such as nhmmer, web services. Pfam ([Link] is a database of protein
cross_match, AB-BLAST/WU-BLAST, RMBLAST, and Decy- families, which contains many protein families, each of which is
pher, and to report all locations, specific classifications and copy represented by MSAs and HMMs. REXdb ([Link]
number information of the hit sequences. The principle of org/?page_id=918) is a reference database of TE protein domains
repetitive DNA sequence classification and masking is presented employed in the repeat analysis tools RepeatExplorer217 and
in Fig. 5. DANTE207, which are available on the Galaxy server (https://
Repeats
ĂĂ A-TT-GTG-GGA-C ĂĂ ĂĂ A-TT-GTG-GGAGC ĂĂ
Sequence fragment-1 Sequence fragment-3
˅Repe
˄d˅ pe at
at mas
asmki
king
Reference genome / Assemblies / Sequencing reads Repeve sequence nucleic acid libraries Repeve sequence protein libraries
NNNNNNNNNNN
A-TT-GTG-GGA-C ĂĂ Masking
AATTTGTGCGGAGC REXdb
Reports RepBase Dfam msRepDB RepeatsDB Pfam
Sequence fragment-1
NNNNNNNNNNN
ĂĂ A-TT-GTG-GGAGC ĂĂ
AATTTGTGCGGAGC User lib
Sequence fragment-2 SINE/Alu : AATTTGTGCGGAGC
Repeat consensus sequence with classificaon informaon
NNNNNNNNNNN NNNNNNNNNNN
ĂĂ AATT-GTG-GGA-C ĂĂ A-TT-GTGCGGA-C ĂĂ RepeatMasker
AATTTGTGCGGAGC AATTTGTGCGGAGC
Sequence fragment-3
Blast-based aligner
Repeat masking
Fig. 5 The principle of automatic repeat classification and masking. Sub-graph (a): A simple example of the distribution characteristics of repeats in the
reference genome, where the black blocks represent chromosomes. Sub-graph (b): Principle of repeat detection, where the final sequence composed of
colored bases represents the consensus repeat sequence. Sub-graph (c): Principle of automatic repeat classification, where black and dark green cylinders
represent nucleic acid and protein libraries, respectively. Sub-graph (d): Principle of automatic repeat masking. Light green cylinders in Sub-graphs (c) and
(d) represent the user-defined repeat library, and black blocks in Sub-graph (d) indicate the sequencing reads from various samples of the same or similar
species.
[Link]/). A detailed introduction to repeat classify them. Compared with TEclass, TERL can distinguish
databases is shown in Supplementary Note 7.3.2. A performance non-TE sequences and label numerously of unknown types of
comparison of the databases is presented in Supplementary repetitive sequences in the detection results as corresponding
Tables S33–S41. non-TE types, which greatly improves the accuracy of non-TE
sequence identification. In addition, TERL has excellent scal-
ability and can be executed seamlessly in GPUs, greatly
Automated repeat classification methods based on homology
improving the efficiency of data processing.
searching. The goal of classification is to classify unknown
repeats into their main taxonomic branches, which usually refers
to the classification of TEs (Fig. 5(a), (b) and (c)). Some methods Automated masking of repeats. Repeat masking is also a vital step
are proposed based on manually predefined features for auto- in the pipeline of genome repeat analysis (Fig. 5(D)). Three steps of
matically classifying TEs, such as TEclass208, RepeatClassifier206, detection, classification, and masking are integrated into some
PASTEC209, and REPCLASS210. Homology-based searching and hybrid repeat detection frameworks, such as RepeatMasker,
structural features of TEs (e.g., TSD, TRs, tRNA, poly-A signals, RepeatModeler, and LongRepMarker, to obtain classified TEs (e.g.,
SSR, and protein-coding domains) are used in these tools to LTRs, LINEs, SINEs, etc.) and masking reports (e.g., the length
perform classification (Table 5). occupied, coverage ratio, and location of each TE in the genome).
For instance, TEclass ([Link] As described, RepeatMasker ([Link] is a
teclass) uses support vector machine (SVM) and oligomer robust detection and masking framework based on homology
frequencies to classify TE consensus repeat sequences into DNA searching. The input of RepeatMasker are the genome to be
transposons and retrotransposons, including LTRs, LINEs, and annotated and a standard repeat library, such as the RepBase or
SINEs. RepeatClassifier ([Link] Dfam. During the masking process, RepeatMasker aligns the well-
RepeatModeler) is a homology-based classification module classified TEs collected in the repeat library to the sequences of the
designed in the hybrid TE family discovery framework RepeatMo- genome one by one, records the length occupied, coverage ratio,
deler2, which compares TE families to RepeatMasker repeat and location of each TE in the genome, and generates a masking
protein databases (e.g., Pfam, REXdb) and RepeatMasker repeat report. Performance analyzes of automated repeat sequence clas-
nucleic acid libraries (e.g., RepBase and Dfam) using the homology- sification and masking methods are shown in Supplementary
based aligner BLAST. PASTEC ([Link] Tables S42–S46.
PASTEClassifier) obtains the similarities and structural features of
TEs using profile HMMs211 and homology-based search algo-
Discussion
rithms (e.g., tblastx, blastx, and blastn) and then classifies TEs into
In this section, we summarize the challenges and solutions in the
their respective order. REPCLASS ([Link]
research field of genomic repeat detection and annotation, as well
repclass/) is a tool that automates the classification of TE sequences
as future development trends.
using control repeat libraries and structural and homology
Since not requiring prior knowledge, the de novo methods are
characterization modules, which can classify accurately virtually
more flexible and valuable than the homology-based and
any known TR types.
structure-based methods. However, developing advanced de novo
algorithms for comprehensive repetitive DNA sequence detection
Automatic repeat classification methods based on machine and is challenging due to the short length of NGS reads and the high
deep learning. Convolutional neural networks (CNNs) are rate of sequencing errors in TGS (Third-generation sequencing)
automatic and adaptive representation learning and feature reads. A hybrid strategy combining short and long reads is cur-
extraction algorithms that can be applied to predict unknown rently the most effective way to achieve the above goals. However,
sequence profiles or motifs and functional activity discovery before implementing the hybrid strategy, we need to obtain
without pre-defining sequence features. Some TE classification multiple sequencing data, such as NGS reads, TGS reads, and
algorithms are proposed based on CNNs, among which even 10× genomic reads, of the same sample in advance, resulting
DeepTE212 and TERL213 are representatives (Table 5). high detection costs and difficult algorithm design. Therefore,
DeepTE ([Link] tra- nsforms successfully overcoming the impact of sequencing errors in TGS
sequences into input vectors through a k-mer counting strategy, reads and directly carrying out high-precision and ultra-complete
and classifies TEs into superfamilies and orders based on a tree- repeat detection using the increasing number of high-quality TGS
structured classification process and eight trained models (class reads will become a research focus in the future. Furthermore, the
model, classI model, LTR model, nLTR model, SINE model, LINE variation of TRs is closely related to the emergence of complex
model, classII_sub1 model and domain model). Among these diseases, such as cancers, neurological disorders, and autism.
models, class model is responsible for classifying TEs into Class I, However, there has not been much progress in the development
Class II_sub1 and Class II_sub2 transposons, and “ClassI model” of algorithms for the detection of TRs and their expansions.
is to classify TEs into LTR and non-LTR transposons. Moreover, Databases containing TRs of multiple species are also very scarce.
the false classification correction model and distinction algorithm Therefore, researching superior identification methods for TRs
for distinguishing non-TEs and TEs are also integrated into and complete TR databases is of great significance in exploring
DeepTE. TERL ([Link] is a their biological functions in genomes, which is another important
fast and flexible deep CNN-based approach for classifying TEs research focus in the future.
and other biological sequences, which employs deep CNNs to Several automatic repeat classification methods have been
preprocess and translate one-dimensional nucleic acid sequences proposed based on machine and deep Learning. These methods
(i.e., image-like data of nucleic acid sequences) into two- all benefit from SVM and CNNs and perform better than tradi-
dimensional space data. TEclass is an automated classification tional methods in some aspects. However, the completeness of the
algorithm based on machine learning support vector machine classification is very limited. For example, TEclass can only
(SVM). The classification obtained using TEclass is very sparse classify TEs into the following four classes: DNA transposons,
relative to the overall TE classes, usually only including DNA LTR, LINE, and SINE, and its classification results tend to have
transposons, LTRs, LINEs, and SINEs. Besides, TEclass can only high false-positive rates. Moreover, DeepTE uses CNNs to classify
roughly distinguish non-TE sequences, but cannot accurately unknown TEs by converting sequences into input vectors based
on k-mer counting, which can be used to distinguish TEs and Received: 29 March 2023; Accepted: 4 September 2023;
non-TEs with relatively low false-positive rates. Both TEclass and
REPCLASS cannot distinguish between TEs and other non-TEs,
so DeepTE is superior to them. Nevertheless, DeepTE is also not
perfect. First, the completeness of its classification remains
unsatisfactory. Second, DeepTE is not specifically designed to References
classify nested TE, and the databases it depends on do not include 1. Biscotti, M. A., Olmo, E. & Heslop-Harrison, J. S. Repetitive DNA in
annotations for nested TEs. Deep neural networks (DNNs) have eukaryotic genomes. Chromosom. Res. 23, 415–420 (2015).
great application potential in automated repeat classification. 2. Mrázek, J., Guo, X. & Shah, A. Simple sequence repeats in prokaryotic
genomes. Proc. Natl Acad. Sci. USA. 104, 8472–8477 (2007).
However, current methods did not maximize the advantages of
3. Jurka, J., Kapitonov, V. V., Kohany, O. & Jurka, M. V. Repetitive sequences in
DNNs. Therefore, developing superior DNNs and models for complex genomes: structure and evolution. Annu. Rev. Genom. Hum. Genet. 8,
more comprehensive and accurate repeat classification is one of 241–259 (2007).
the main research focuses for the future. 4. Treangen, T. J., Abraham, A. L., Touchon, M. & Rocha, E. P. Genesis, effects
TEs carry cis-regulatory sequences that can alter gene reg- and fates of repeats in prokaryotic genomes. FEMS Microbiol. Rev. 33,
539–571 (2009).
ulatory networks through redistributing transcription factor 5. Bernabe, I. B. et al. Genome-wide contribution of common short-tandem
binding sites and developing novel enhancer activities. Its repeats to Parkinson’s disease genetic risk. Brain 146, 65–74 (2023).
abnormal expression is closely related to many complex diseases, 6. Nekrutenko, A. & Li, W. H. Transposable elements are found in a large
such as cancers. However, the role of TEs in cell-type hetero- number of human protein-coding genes. Trends Genet. 17, 619–621
geneity and biological processes has not been fully revealed, and (2001).
7. Alexander, R. P., Fang, G., Rozowsky, J., Snyder, M. & Gerstein, M. B.
research in this field is still in its infancy. With the rapid devel-
Annotating non-coding regions of the genome. Nat. Rev. Genet. 11, 559–571
opment of single-cell technologies, scRNA-seq has become an (2010).
efficient method for observing cell activity, which can be used to 8. Bourque, G. et al. Ten things you should know about transposable elements.
analyze gene-centric and TE expression accurately. Therefore, a Genome Biol. 19, 199 (2018).
future research focus is to quantify TE expression and explore the 9. Zhang, X. & Meyerson, M. Illuminating the noncoding genome in cancer. Nat.
Cancer 1, 864–872 (2020).
role of TEs in the pathway and mechanism of complex diseases at 10. Mehrotra, S. & Goyal, V. Repetitive Sequences in Plant Nuclear DNA: Types,
the single-cell level. Distribution, Evolution and Function. Genom. Proteom. Bioinform. 12,
164–171 (2014).
Conclusion 11. Zu, T. et al. Non-ATG-initiated translation directed by microsatellite
expansions. Proc. Natl Acad. Sci. USA. 108, 260–5 (2011).
Repetitive DNA sequences play an indispensable role in the
12. Al-Turki, T. M. & Griffith, J. D. Mammalian telomeric RNA (TERRA) can be
physiological activities of organisms, and they comprise almost translated to produce valine-arginine and glycine-leucine dipeptide repeat
half of the human genome. Repeats in genomes can be divided proteins. Proc. Natl Acad. Sci. USA. 120, e2221529120 (2023).
into TEs and TRs. TEs can result in mutations, altered gene 13. Hannan, A. J. Tandem repeats mediating genetic plasticity in health and
expression, chromosome rearrangement. etc., which are related to disease. Nat. Rev. Genet. 19, 286–298 (2018).
14. Ishiura, H. et al. Noncoding CGG repeat expansions in neuronal intranuclear
many diseases, such as cancers, genetic disorders, autoimmune
inclusion disease, oculopharyngodistal myopathy and an overlapping disease.
diseases, and metabolic disorders. TRs, especially STRs, are highly Nat. Genet. 51, 1222–1232 (2019).
variable, which can accelerate the gene expression evolution and 15. Shah, N. M. et al. Pan-cancer analysis identifies tumor-specific antigens
generate sufficient variability that allows a rapid evolution and derived from transposable elements. Nat. Genet. 55, 631–639 (2023). This
adaptation of organisms, and play a vital role in the structural article reported that cryptic promoters within transposable elements (TEs) can
stability of genetic materials and regulate gene expression, causing be transcriptionally reactivated in tumors to create new TE-chimeric
transcripts, which can produce immunogenic antigens.
various disorders. Due to a lack of sufficiently advanced detection 16. Touati, R. et al. New methodology for repetitive sequences identification in
technologies, the role and effect of repeats in genomes, especially human X and Y chromosomes. Biomed. Signal Proc. Control 64, 102207
the human genome, have been underestimated. We believe that (2021).
this review will be helpful in the understanding of repeats in 17. Novák, P. et al. TAREAN: a computational tool for identification and
characterization of satellite DNA from unassembled short reads. Nucleic Acids
genomes and provide guidance for repeat annotation (detection,
Res. 45, e111–e111 (2017).
classification, and masking) and in-depth exploration of its 18. Liehr, T. Repetitive elements in humans. Int. J. Mol. Sci. 22, 2072 (2021).
association with human diseases. 19. Novák, P., Neumann, P. & Macas, J. Global analysis of repetitive DNA from
unassembled sequence reads using RepeatExplorer2. Nat. Protoc. 15,
Reporting summary. Further information on research design is 3745–3776 (2020).
20. McNulty, S. M. & Sullivan, B. A. Alpha satellite DNA biology: finding function
available in the Nature Portfolio Reporting Summary linked to in the recesses of the genome. Chromosom. Res. 26, 115–138 (2018).
this article. 21. Youssef, N., Budd, A. & Bielawski, J. P. Introduction to Genome Biology and
Diversity. Methods Mol. Biol. 1910, 3–31 (2019).
22. Bishop, C. E., Guellaen, G., Geldwerth, D. VossR., Fellous, M. & Weissenbach,
J. Single-copy DNA sequences specific for the human Y chromosome. Nature
Data availability 303, 831–832 (1983).
The reference genomes of six species: Homo sapiens (GCF_000001405.39), Gallus
23. Hou, Z., Romero, R., Uddin, M., Than, N. G. & Wildman, D. E. Adaptive
(GCF_016699485.2), Mouse (GCF_000001635.27), Drosophila melanogaster
history of single copy genes highly expressed in the term human placenta.
(GCA_018903765.1), Glycine max (GCA_000004515.5) and Leafcutter ant
Genomics 93, 33–41 (2009).
(GCA_000204515.1) are downloaded from the NCBI website ([Link]
24. Pavlicek A., Kapitonov V.V., & Jurka J. Human Repetitive DNA[M].
gov/). Five groups of NGS short reads: Leafcutter Ant (ERR034186, [Link]
Encyclopedic Reference of Genomics and Proteomics in Molecular Medicine.
[Link]/), [Link] (SRR350 908, [Link] Mouse
(Springer, Berlin, Heidelberg, 2005).
(ERR2894257, [Link] Human-chr14([Link]
25. Kojima, K. K. Structural and sequence diversity of eukaryotic transposable
edu/) and HG003_24149_father (D2 S2 L001 R1 001, [Link]
elements. Genes Genet. Syst. 94, 233–252 (2020).
giab/ftp/data), three groups of barcode linked reads (HG003_24149_father,
26. Genovese, L. M. et al. A Census of Tandemly Repeated Polymorphic Loci in
HG004_NA24143, and HG002_NA24385_son, [Link]
Genic Regions Through the Comparative Integration of Human Genome
data), three groups of CCS long reads (HG003_24149_father, HG004_NA24143_mother
Assemblies. Front. Genet. 9, 155 (2018).
and HG002_NA24385_son, [Link] and four
27. Richard, G. F., Kerrest, A. & Dujon, B. Comparative genomics and molecular
groups of PacBio long reads (dro_100k, human_100k, dmel_filtered and
dynamics of DNA repeats in eukaryotes. Microbiol. Mol. Biol. Rev. 72,
human_polished, [Link] are used to evaluate the
686–727 (2008).
performance of each tool in this study.
28. Sullivan, L. L., Chew, K. & Sullivan, B. A. α satellite DNA variation and 59. Lerat, E. & Capy, P. Retrotransposons and retroviruses: analysis of the
function of the human centromere. Nucleus 8, 331–339 (2017). envelope gene. Mol. Biol. Evol. 16, 1198–1207 (1999).
29. Sawaya, S. et al. Microsatellite tandem repeats are abundant in human 60. Havecker, E. R., Gao, X. & Voytas, D. F. The diversity of LTR
promoters and are associated with regulatory elements. PLoS ONE 8, e54710 retrotransposons. Genome Biol. 5, 225 (2004).
(2013). 61. Gro˙ger, V. et al. Formation of HERV-K and HERV-Fc1 Envelope Family
30. Richard, G. F. & Pâques, F. Mini- and microsatellite expansions: the Members is Suppressed on Transcriptional and Translational Level. Int. J.
recombination connection. EMBO Rep. 1, 122–126 (2000). Mol. Sci. 21, 7855 (2020).
31. Li, H. Identifying centromeric satellites with dna-brnn. Bioinformatics 35, 62. Nelson, P. N. et al. Human endogenous retroviruses: transposable elements
4408–4410 (2019). with potential? Clin. Exp. Immunol. 138, 1–9 (2004).
32. Alaguponniah, S. et al. Finding of novel telomeric repeats and their 63. Zhao, J. et al. Expression of Human Endogenous Retrovirus Type K Envelope
distribution in the human genome. Genomics 112, 3565–3570 (2020). Protein is a Novel Candidate Prognostic Marker for Human Breast Cancer.
33. Riethman, H. Human subtelomeric copy number variations. Cytogenet. Genes Cancer 2, 914–922 (2011).
Genome Res. 123, 244–252 (2008). 64. Sohn, J. & Nam, J. W. The present and future of de novo whole-genome
34. Han, Y. & Wessler, S. R. MITE-Hunter: a program for discovering miniature assembly. Brief Bioinform. 19, 23–40 (2018).
inverted-repeat transposable elements from genomic sequences. Nucleic Acids 65. Liao, X. et al. Current challenges and solutions of de novo assembly. Quant.
Res. 38, e199 (2010). Biol. 7, 90–109 (2019).
35. Kojima, K. K. & Jurka, J. Crypton transposons: identification of new diverse 66. Kamath, G. M. et al. HINGE: long-read assembly achieves optimal repeat
families and ancient domestication events. Mobile DNA 2, 12 (2011). resolution. Genome Res. 27, 747–756 (2017). This article reported an
36. Krupovic, M. & Koonin, E. V. Polintons: a hotbed of eukaryotic virus, assembler that seeks to achieve optimal repeat resolution by distinguishing
transposon and plasmid evolution. Nat. Rev. Microbiol. 13, 105–115 (2015). repeats that can be resolved given the data from those that cannot.
37. Lee, T. F. et al. RNA polymerase V-dependent small RNAs in Arabidopsis 67. Jain, C. et al. Long-read mapping to repetitive reference sequences using
originate from small, intergenic loci including most SINE repeats. Epigenetics Winnowmap2. Nat. Methods 19, 705–710 (2022).
7, 781–795 (2012). 68. Jakubosky, D. et al. Properties of structural variants and short tandem repeats
38. Pace, J. K. & Feschotte, C. The evolutionary history of human DNA associated with gene expression and complex traits. Nat. Commun. 11, 2927
transposons: evidence for intense activity in the primate lineage. Genome Res. (2020).
17, 422–432 (2007). 69. Liao, X. et al. Improving de novo assembly based on read classification. IEEE/
39. Muñoz-López, M. & Garcĺa-Pérez, J. L. DNA transposons: nature and ACM Trans. Comput. Biol. Bioinform. 17, 177–188 (2018).
applications in genomics. Curr. Genom. 11, 115–128 (2010). 70. Miga, K. H. et al. Telomere-to-telomere assembly of a complete human X
40. Kojima, K. K. Human transposable elements in Repbase: genomic footprints chromosome. Nature 585, 79–84 (2020).
from fish to humans. Mobile DNA 9, 2 (2018). 71. Narzisi, G. & Schatz, M. C. The challenge of small-scale repeats for indel
41. David, J. F. Retrotransposons. Curr. Biol. 22, R432–R437 (2012). discovery. Front. Bioeng. Biotechnol. 3, 8 (2015).
42. Muszewska, A., Hoffman-Sommer, M. & Grynberg, M. LTR retrotransposons 72. Trigiante, G., Blanes, R. N. & Cerase, A. Emerging Roles of Repetitive and
in fungi. PLoS ONE 6, e29425 (2011). Repeat-Containing RNA in Nuclear and Chromatin Organization and Gene
43. Thompson, P. J., Macfarlan, T. S. & Lorincz, M. C. Long Terminal Repeats: Expression. Front. Cell Dev. Biol. 9, 735527 (2021).
From Parasitic Elements to Building Blocks of the Transcriptional Regulatory 73. Gao, D. et al. Transposons play an important role in the evolution and
Repertoire. Mol. Cell 62, 766–76 (2016). diversification of centromeres among closely related species. Front. Plant Sci.
44. Ardeljan, D., Taylor, M. S., Ting, D. T. & Burns, K. H. The Human Long 6, 216 (2015).
Interspersed Element-1 Retrotransposon: An Emerging Biomarker of 74. Nishihara, H. Transposable elements as genetic accelerators of evolution:
Neoplasia. Clin. Chem. 63, 816–822 (2017). contribution to genome size, gene regulatory network rewiring and
45. Kramerov, D. A. & Vassetzky, N. S. Origin and evolution of SINEs in morphological innovation. Genes Genet. Syst. 94, 269–281 (2020).
eukaryotic genomes. Heredity 107, 487–495 (2011). 75. Ramakrishnan, M. et al. The Dynamism of Transposon Methylation for Plant
46. Han, G. et al. Diversity of short interspersed nuclear elements (SINEs) in Development and Stress Adaptation. Int. J. Mol. Sci. 22, 11387 (2021).
lepidopteran insects and evidence of horizontal SINE transfer between 76. Chuong, E. B., Elde, N. C. & Feschotte, C. Regulatory activities of transposable
baculovirus and lepidopteran hosts. BMC Genom. 22, 226 (2021). elements: from conflicts to benefits. Nat. Rev. Genet. 18, 71–86 (2017).
47. Malicki, M., Spaller, T., Winckler, T. & Hammann, C. DIRS retrotransposons 77. González, J. et al. High rate of recent transposable element-induced adaptation
amplify via linear, single-stranded cDNA intermediates. Nucleic Acids Res. 48, in Drosophila melanogaster. PLoS Biol. 6, e251 (2008).
4230–4243 (2020). 78. Ayarpadikannan, S. & Kim, H. S. The impact of transposable elements in
48. Wiegand, S. et al. The Dictyostelium discoideum RNA-dependent RNA genome evolution and genetic instability and their implications in various
polymerase RrpC silences the centromeric retrotransposon DIRS-1 post- diseases. Genom. Inform. 12, 98–104 (2014).
transcriptionally and is required for the spreading of RNA silencing signals. 79. Hancks, D. C. & Kazazian, H. H. Roles for retrotransposon insertions in
Nucleic Acids Res. 42, 3330–3345 (2014). human disease. Mobile DNA 7, 9 (2016).
49. Wang, Y., Gallagher-Jones, M., Suśac, L., Song, H. & Feigon, J. A structurally 80. Voronova, A. et al. Retrotransposon distribution and copy number variation
conserved human and Tetrahymena telomerase catalytic core. Proc. Natl in gymnosperm genomes. Tree Genet. Genomes 13, 88 (2017).
Acad. Sci. USA. 117, 31078–31087 (2020). 81. Pavlicek, A., Gentles, A. J., Paces, J., Paces, V. & Jurka, J. Retroposition of
50. Arkhipova, I. R. Distribution and Phylogeny of Penelope-Like Elements in processed pseudogenes: the impact of RNA stability and translational control.
Eukaryotes. Syst. Biol. 55, 875–885 (2006). Trends Genet. 22, 69–73 (2006).
51. Gladyshev, E. A. & Arkhipova, I. R. Telomere-associated endonuclease- 82. Ovchinnikov, I., Troxel, A. B. & Swergold, G. D. Genomic characterization of
deficient Penelope-like retroelements in diverse eukaryotes. Proc. Natl Acad. recent human LINE-1 insertions: evidence supporting random insertion.
Sci. USA. 104, 9352–9357 (2007). Genome Res. 11, 2050–2058 (2001).
52. Han, J. S. Non-long terminal repeat (non-LTR) retrotransposons: 83. Ponomaryova, A. A. et al. Aberrant Methylation of LINE-1 Transposable
mechanisms, recent developments, and unanswered questions. Mobile DNA 1, Elements: A Search for Cancer Biomarkers. Cells 9, 2017 (2020).
15 (2010). 84. McKerrow, W. et al. LINE-1 expression in cancer correlates with p53
53. Scott, E. C. et al. A hot L1 retrotransposon evades somatic repression and mutation, copy number alteration, and S phase checkpoint. Proc. Natl Acad.
initiates human colorectal cancer. Genome Res. 26, 745–755 (2016). Sci. USA. 119, e2115999119 (2022). This article reported that LINE-1
54. Miki, Y. et al. Disruption of the APC gene by a retrotransposal insertion of expression in cancer correlates with p53 mutation, copy number alteration,
L1 sequence in a colon cancer. Cancer Res. 52, 643–645 (1992). and S phase checkpoint.
55. Larsen, P. A. et al. The Alu neurodegeneration hypothesis: A primate-specific 85. Witherspoon, D. J. et al. Mobile element scanning (ME-Scan) identifies
mechanism for neuronal transcription noise, mitochondrial dysfunction, and thousands of novel Alu insertions in diverse human populations. Genome Res.
manifestation of neurodegenerative disease. Alzheimers Dement. 13, 828–838 23, 107–116 (2013).
(2017). 86. Savage, A. L. et al. Characterisation of retrotransposon insertion
56. Payer, L. M. et al. Structural variants caused by Alu insertions are associated polymorphisms in whole genome sequencing data from individuals with
with risks for many human diseases. Proc. Natl Acad. Sci. USA. 114, amyotrophic lateral sclerosis. Gene 843, 146799 (2022).
E3984–E3992 (2017). 87. Zhang, Y. et al. Transcriptionally active HERV-H retrotransposons demarcate
57. Gianfrancesco, O., Bubb, V. J. & Quinn, J. P. SVA retrotransposons as potential topologically associating domains in human pluripotent stem cells. Nat. Genet.
modulators of neuropeptide gene expression. Neuropeptides 64, 3–7 (2017). 51, 1380–1388 (2019).
58. Petrozziello, T. et al. SVA insertion in X-linked Dystonia Parkinsonism alters 88. Uzunović, J., Josephs, E. B., Stinchcombe, J. R. & Wright, S. I. Transposable
histone H3 acetylation associated with TAF1 gene. PLoS ONE 15, e0243655 Elements Are Important Contributors to Standing Variation in Gene
(2020). Expression in Capsella Grandiflora. Mol. Biol. Evol. 36, 1734–1745 (2019).
89. Chishima, T., Iwakiri, J. & Hamada, M. Identification of Transposable 118. Fueyo, R., Judd, J., Feschotte, C. & Wysocka, J. Roles of transposable elements
Elements Contributing to Tissue-Specific Expression of Long Non-Coding in the regulation of mammalian transcription. Nat. Rev. Mol. Cell Biol. 23,
RNAs. Genes 9, 23 (2018). 481–497 (2022). This article reported that TEs often contain sequences
90. Horváth, V., Merenciano, M. & González, J. Revisiting the Relationship capable of recruiting the host transcription machinery, which they use to
between Transposable Elements and the Eukaryotic Stress Response. Trends express their own products and promote transposition.
Genet. 33, 832–841 (2017). 119. Usdin, K. The biological effects of simple tandem repeats: lessons from the
91. Anastasia, A. Z. et al. Transcriptional regulation of human-specific SVAF1 repeat expansion diseases. Genome Res. 18, 1011–1019 (2008).
retrotransposons by cis-regulatory MAST2 sequences. Gene 505, 128–136 120. Haubold, B. & Wiehe, T. How repetitive are genomes? BMC Bioinform. 7,
(2012). 541–551 (2006).
92. Barnada, S. M. et al. Genomic features underlie the co-option of SVA 121. Yi, H. et al. The Tandem Repeats Enabling Reversible Switching between the
transposons as cis-regulatory elements in human pluripotent stem cells. PLoS Two Phases of β-Lactamase Substrate Spectrum. PLOS Genet. 10, e1004640
Genet. 18, e1010225 (2022). (2014).
93. Zhang, X. O., Gingeras, T. R. & Weng, Z. Genome-wide analysis of 122. Bulik-Sullivan, B. et al. An atlas of genetic correlations across human diseases
polymerase III-transcribed Alu elements suggests cell-type-specific enhancer and traits. Nat. Genet. 47, 1236–1241 (2015).
function. Genome Res. 29, 1402–1414 (2019). 123. O’Dushlaine, C. T., Edwards, R. J., Park, S. D. & Shields, D. C. Tandem repeat
94. Lupski, J. R. & Stankiewicz, P. Genomic disorders: molecular mechanisms for copy-number variation in protein-coding regions of human genes. Genome
rearrangements and conveyed phenotypes. PLoS Genet. 1, e49 (2005). Biol. 6, R69 (2005).
95. Cordaux, R. & Batzer, M. A. The impact of retrotransposons on human 124. Hannan, A. J. Tandem repeat polymorphisms: Mediators of genetic plasticity
genome evolution. Nat. Rev. Genet. 10, 691–703 (2009). modulators of biological diversity and dynamic sources of disease
96. Klein, S. J. & O’Neill, R. J. Transposable elements: genome innovation, susceptibility. Adv. Exp. Med. Biol. 769, 1–9 (2012).
chromosome diversity, and centromere conflict. Chromosom. Res. 26, 5–23 125. Fan, H. & Chu, J. Y. A brief review of short tandem repeat mutation. Genom.
(2018). Proteom. Bioinform. 5, 7–14 (2007).
97. Burns, K. Transposable elements in cancer. Nat. Rev. Cancer 17, 415–424 126. Castillo-Lizardo, M., Henneke, G. & Viguera, E. Replication slippage of the
(2017). This article reported that the activity of transposable elements in thermophilic DNA polymerases B and D from the Euryarchaeota Pyrococcus
human cancers, particularly long interspersed element-1 (LINE-1), leads to abyssi. Front. Microbiol. 5, 403 (2014).
somatically acquired insertions in cancer genomes. 127. Gymrek, M., Willems, T., Reich, D. & Erlich, Y. Interpreting short tandem
98. Ahmadi, A. et al. Transposable elements in brain health and disease. Ageing repeat variations in humans using mutational constraint. Nat. Genet. 49,
Res. Rev. 64, 101153 (2020). This article reported that TEs are expressed and 1495–1501 (2017).
active in the brain, challenging the dogma that neuronal genomes are static 128. Gemayel, R., Vinces, M. D., Legendre, M. & Verstrepen, K. J. Variable tandem
and revealing that they are susceptible to somatic genomic alterations, and repeats accelerate evolution of coding and regulatory sequences. Annu. Rev.
have a role in behavior and cognition. Genet. 44, 445–477 (2010).
99. Saleh, A., Macia, A. & Muotri, A. R. Transposable Elements, Inflammation, 129. Mukamel, R. E. et al. Protein-coding repeat polymorphisms strongly shape
and Neurological Disease. Front. Neurol. 10, 894 (2019). diverse human phenotypes. Science 373, 1499–1505 (2021).
100. Kim, Y. J., Lee, J. & Han, K. Transposable Elements: No More ’Junk DNA’. 130. Farré, M., Bosch, M., López-Giráldez, F., Ponsá, M. & Ruiz-Herrera, A.
Genom. Inform. 10, 226–233 (2012). Assessing the role of tandem repeats in shaping the genomic architecture of
101. Balachandran, P. et al. Transposable element-mediated rearrangements are great apes. PLoS ONE 6, e27239 (2011).
prevalent in human genomes. Nat. Commun. 13, 7115 (2022). This article 131. Gemayel, R., Cho, J., Boeynaems, S. & Verstrepen, K. J. Beyond junk-variable
reported that the transposable element-mediated rearrangements are enriched tandem repeats as facilitators of rapid evolution of regulatory and coding
in genic loci and can create potentially important risk alleles such as a deletion sequences. Genes 3, 461–80 (2012).
in TRIM65, a known cancer biomarker and therapeutic target. 132. Shi, Y. et al. Characterization of genome-wide STR variation in 6487 human
102. Niu, Y. et al. Characterizing mobile element insertions in 5675 genomes. genomes. Nat. Commun. 14, 2092 (2023). This article reported that short
Nucleic Acids Res. 50, 2493–2508 (2022). tandem repeat mutations were affected by motif length, chromosome context
103. Huang, C. R., Burns, K. H. & Boeke, J. D. Active transposition in genomes. and epigenetic features.
Annu. Rev. Genet. 46, 651–675 (2012). 133. Fotsing, S. F. et al. The impact of short tandem repeat variation on gene
104. Cordaux, R., Hedges, D. J., Herke, S. W. & Batzer, M. A. Estimating the expression. Nat. Genet. 51, 1652–1659 (2019). This article reported that
retrotransposition rate of human Alu elements. Gene 373, 134–137 (2006). expression of short tandem repeats explain a sizable portion (10–15%) of the
105. Rosser, J. M. & An, W. L1 expression and regulation in humans and rodents. cis heritability of gene expression.
Front. Biosci. (Elite Ed) 4, 2203–2225 (2012). 134. Aguilar, M. & Prieto, P. Telomeres and Subtelomeres Dynamics in the Context
106. Chuang, N. T. et al. Mutagenesis of human genomes by endogenous mobile of Early Chromosome Interactions During Meiosis and Their Implications in
elements on a population scale. Genome Res. 31, 2225–35 (2021). Plant Breeding. Front. Plant Sci. 12, 672489 (2021).
107. Payer, L. M. & Burns, K. H. Transposable elements in human genetic disease. 135. Lamb, J. C. & Birchler, J. A. The role of DNA sequence in centromere
Nat. Rev. Genet. 20, 760–772 (2019). This article reviewed many ways human formation. Genome Biol. 4, 214 (2003).
retrotransposons contribute to genome function, their dysregulation in 136. Miga, K. H. & Alexandrov, I. A. Variation and evolution of human
diseases including cancer, and how they affect genetic disease. centromeres: a field guide and perspective. Ann. Rev. Genet. 55, 583–602
108. Kannan, S. et al. Transposable Element Insertions in Long Intergenic Non- (2021).
Coding RNA Genes. Front. Bioeng. Biotechnol. 3, 71 (2015). 137. Lim, C. J. & Cech, T. R. Shaping human telomeres: from shelterin and CST
109. Etchegaray, E., Naville, M., Volff, J. N. & Haftek-Terreau, Z. Transposable complexes to telomeric chromatin organization. Nat. Rev. Mol. Cell Biol. 22,
element-derived sequences in vertebrate development. Mob. DNA 12, 1 283–298 (2021).
(2021). 138. Sun, J. H. et al. Disease-Associated Short Tandem Repeats Co-localize with
110. Johnson, R. & Guigó, R. The RIDL hypothesis: transposable elements as Chromatin Domain Boundaries. Cell 175, 224-238.e15 (2018).
functional domains of long noncoding RNAs. RNA 20, 959–976 (2014). 139. Ishiura, H. et al. Expansions of intronic TTTCA and TTTTA repeats in benign
111. Cuevas-Diaz, D. R. et al. Long non-coding RNAs: important regulators in the adult familial myoclonic epilepsy. Nat. Genet. 50, 581–590 (2018).
development, function and disorders of the central nervous system. 140. Albertin, C. B. et al. The octopus genome and the evolution of cephalopod
Neuropathol. Appl. Neurobiol. 45, 538–556 (2019). neural and morphological novelties. Nature 524, 220–224 (2015).
112. Grandi, N. & Tramontano, E. HERV Envelope Proteins: Physiological Role 141. DeJesus-Hernandez, M. et al. Expanded GGGGCC hexanucleotide repeat in
and Pathogenic Potential in Cancer and Autoimmunity. Front. Microbiol. 9, noncoding region of C9orf72 causes chromosome 9p-linked FTD and ALS.
462 (2018). Neuron 72, 245–256 (2011).
113. Mao, J., Zhang, Q. & Cong, Y. S. Human endogenous retroviruses in 142. Duan, Y. et al. PARylation regulates stress granule dynamics, phase separation,
development and disease. Comput. Struct. Biotechnol. J. 19, 5978–5986 (2021). and neurotoxicity of disease-related RNA-binding proteins. Cell Res. 29,
114. Hermant, C. & Torres-Padilla, M. E. TFs for TEs: the transcription factor 233–247 (2019).
repertoire of mammalian transposable elements. Genes Dev. 35, 22–39 (2021). 143. Raghupathy, N. & Durand, D. Gene cluster statistics with gene families. Mol.
115. Senft, A. D. & Macfarlan, T. S. Transposable elements shape the evolution of Biol. Evol. 26, 957–968 (2009).
mammalian development. Nat. Rev. Genet. 22, 691–711 (2021). 144. Bonthala, V. S. & Stich, B. Genetic Divergence of Lineage-Specific Tandemly
116. Evrony, G. D. et al. Single-neuron sequencing analysis of L1 retrotransposition Duplicated Gene Clusters in Four Diploid Potato Genotypes. Front. Plant Sci.
and somatic mutation in the human brain. Cell 151, 483–496 (2012). 13, 875202 (2022).
117. Ali, A., Han, K. & Liang, P. Role of transposable elements in gene regulation in 145. Kuzmin, E., Taylor, J. S. & Boone, C. Retention of duplicated genes in
the human genome. Life 11, 118 (2021). evolution. Trends Genet. 38, 59–72 (2022).
146. Sultanov, D. & Hochwagen, A. Varying strength of selection contributes to the 176. Chen, Y., Zhou, F., Li, G. & Xu, Y. MUST: a system for identification of
intragenomic diversity of rRNA genes. Nat. Commun. 13, 7245 (2022). miniature inverted-repeat transposable elements and applications to
147. Blokhina, Y. P. & Buchwalter, A. Moving fast and breaking things: Incidence Anabaena variabilis and Haloquadratum walsbyi. Gene 436, 1–7 (2009).
and repair of DNA damage within ribosomal DNA repeats. Mutat. Res. 821, 177. Ye, C., Ji, G. & Liang, C. detectMITE: A novel approach to detect miniature
111715 (2020). inverted repeat transposable elements in genomes. Sci. Rep. 6, 19688 (2016).
148. Pajic, P. et al. A mechanism of gene evolution generating mucin function. Sci. 178. Yang, G. MITE Digger, an efficient and accurate algorithm for genome wide
Adv. 8, eabm8757 (2022). discovery of miniature inverted repeat transposable elements. BMC Bioinform.
149. Gymrek, M. et al. Abundant contribution of short tandem repeats to gene 14, 186 (2013).
expression variation in humans. Nat. Genet. 48, 22–29 (2016). 179. Crescente, J. M., Zavallo, D., Helguera, M. & Vanzetti, L. S. MITE Tracker: an
150. Malik, I., Kelley, C. P., Wang, E. T. & Todd, P. K. Molecular mechanisms accurate approach to identify miniature inverted-repeat transposable elements
underlying nucleotide repeat expansion disorders. Nat. Rev. Mol. Cell Biol. 22, in large genomes. BMC Bioinform. 19, 348 (2018).
589–607 (2021). 180. Lerat, E. Identifying repeats and transposable elements in sequenced genomes:
151. Trost, B. et al. Genome-wide detection of tandem DNA repeats that are how to find your way through the dense forest of programs. Heredity 104,
expanded in autism. Nature 586, 80–86 (2020). 520–533 (2010).
152. Chintalaphani, S. R. et al. An update on the neurological short tandem repeat 181. Agarwal, P. & States, D. J. The Repeat Pattern Toolkit (RPT): analyzing the
expansion disorders and the emergence of long-read sequencing diagnostics. structure and evolution of the C. elegans genome. Proc. Int. Conf. Intell. Syst.
Acta Neuropathol. Commun. 9, 98 (2021). Mol. Biol. 2, 1–9 (1994).
153. Depienne, C. & Mandel, J. L. 30 years of repeat expansion disorders: What 182. Chen, G. L., Chang, Y. J. & Hsueh, C. H. PRAP: an ab initio software package
have we learned and what are the remaining challenges? Am. J. Hum. Genet. for automated genome-wide analysis of DNA repeats for prokaryotes.
108, 764–785 (2021). This article reported the development and remaining Bioinformatics 29, 2683–2689 (2013).
challenges of the field of repeat expansion disorders over the past 30 years. 183. Robert, C. E. & Eugene, W. M. PILER: identification and classification of
154. Goodman, L. D. & Bonini, N. M. New Roles for Canonical Transcription genomic repeats. Bioinformatics 21, i152–i158 (2005).
Factors in Repeat Expansion Diseases. Trends Genet. 36, 81–92 (2020). 184. Nicolas, J., Tempel, S., Fiston-Lavier, A. S. & Cherif, E. Finding and
155. Chen, W., Swanson, B. J. & Frankel, W. L. Molecular genetics of characterizing repeats in plant genomes. Methods Mol. Biol. 2443, 327–385
microsatellite-unstable colorectal cancer for pathologists. Diagn. Pathol. 12, 24 (2016).
(2017). 185. Liao, X. et al. A sensitive repeat identification framework based on short and
156. Taylor, J. P., Brown Jr, R. H. & Cleveland, D. W. Decoding ALS: from genes to long reads. Nucleic Acids Res. 49, e100–e100 (2021).
mechanism. Nature 539, 197–206 (2016). 186. Saha, S., Bridges, S., Magbanua, Z. V. & Peterson, D. G. Empirical comparison
157. Bao, W., Kojima, K. K. & Kohany, O. Repbase Update, a database of repetitive of ab initio repeat finding programs. Nucleic Acids Res. 36, 2284–2294 (2008).
elements in eukaryotic genomes. Mobile DNA 6, 11–17 (2015). 187. Price, A. L., Jones, N. C. & Pevzner, P. A. De novo identification of repeat
158. Hubley, R. et al. The Dfam database of repetitive DNA families. Nucleic Acids families in large genomes. Bioinformatics 21, i351–i358 (2005).
Res. 44, D81–D89 (2016). 188. Li, R. et al. ReAS: Recovery of ancestral sequences for transposable elements
159. Liao, X. et al. msRepDB: a comprehensive repetitive sequence database of over from the unassembled reads of a whole genome shotgun. PLoS Comput. Biol.
80 000 species. Nucleic Acids Res. 50, D236–D245 (2021). 1, e43 (2005).
160. Neumann, P. et al. Systematic survey of plant LTR-retrotransposons elucidates 189. Shi, J. & Liang, C. Generic Repeat Finder: A High-Sensitivity Tool for
phylogenetic relationships of their polyprotein domains and provides a Genome-Wide De Novo Repeat Detection. Plant. Physiol. 180, 1803–1815
reference for element classification. Mobile DNA 10, 1–18 (2019). (2019).
161. Jaina, M. et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 190. Koch, P., Platzer, M. & Downie, B. R. RepARK-de novo creation of repeat
49, D412–D419 (2021). libraries from whole-genome NGS reads. Nucleic Acids Res. 42, e80–e80
162. Scott, M. & Thomas, L. M. BLAST: at the core of a powerful and diverse set of (2014).
sequence analysis tools. Nucleic Acids Res. 32, W20–W25 (2004). 191. Chu, C., Nielsen, R. & Wu, Y. REPdenovo: inferring de novo repeat motifs
163. Jurka, J., Klonowski, P., Dagman, V. & Pelton, P. CENSOR-a program for from short sequence reads. PloS ONE 11, e0150719 (2016).
identification and elimination of repetitive elements from DNA sequences. 192. Liao, X., Gao, X., Zhang, X., Wu, F. X. & Wang, J. RepAHR: an improved
Comput. Chem. 20, 119–121 (1996). approach for de novo repeat identification by assembly of the high-frequency
164. Kennedy, R. C. et al. An automated homology-based approach for identifying reads. BMC Bioinform. 21, 463 (2020).
transposable elements. BMC Bioinform. 12, 130 (2011). 193. Guo, R. et al. RepLong: de novo repeat identification using long read
165. Li, X., Kahveci, T. & Settles, A. M. A novel genome-scale repeat finder geared sequencing data. Bioinformatics 34, 1099–1107 (2017).
towards transposons. Bioinformatics 24, 468–476 (2007). 194. Kolpakov, R., Bana, G. & Kucherov, G. mreps: Efficient and flexible detection
166. Fiston-Lavier, A. S., Carrigan, M., Petrov, D. A. & González, J. T-lex: a of tandem repeats in DNA. Nucleic Acids Res. 31, 3672–8 (2003).
program for fast and accurate assessment of transposable element presence 195. Benson, G. Tandem repeats finder: a program to analyze DNA sequences.
using next-generation sequencing data. Nucleic Acids Res. 39, e36 Nucleic Acids Res. 27, 573–80 (1999). This article reported tandem repeat
(2010). finder (TRF), currently the most well-known tandem repeat detection tool.
167. Wicker, T. et al. A unified classification system for eukaryotic transposable 196. Jorda, J. & Kajava, A. V. T-REKS: identification of Tandem REpeats in
elements. Nat. Rev. Genet. 8, 973–982 (2007). sequences with a K-meanS based algorithm. Bioinformatics 25, 2632–8 (2009).
168. Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible 197. Wlodzimierz, P., Hong, M. & Henderson, I. R. TRASH: Tandem Repeat
software for de novo detection of LTR retrotransposons. BMC Bioinform. 9, 18 Annotation and Structural Hierarchy. Bioinformatics 39, btad308 (2023).
(2008). This article reported LTRharvest, currently the most well-known LTR 198. Jam H. Z. et al. A deep population reference panel of tandem repeat variation.
retrotransposon detection tool. bioRxiv 2023.03.09.531600, 1–37 (2023).
169. Darzentas, N., Bousios, A., Apostolidou, V. & Tsaftaris, A. S. MASiVE: 199. Fazal S. et al. RExPRT: a machine learning tool to predict pathogenicity of
Mapping and Analysis of SireVirus Elements in plant genome sequences. tandem repeat loci. bioRxiv 2023.03.22.533484, 1–30 (2023).
Bioinformatics 26, 2452–2454 (2010). 200. Dolzhenko, E. et al. ExpansionHunter: a sequence-graph-based tool to analyze
170. Rho, M., Choi, J. H., Kim, S., Lynch, M. & Tang, H. De novo identification of variation in short tandem repeat regions. Bioinformatics 35, 4754–4756
LTR retrotransposons in eukaryotic genomes. BMC Genom. 8, 90 (2007). (2019).
171. Matej, L., Pavel, J., Ivan, V., Michal, C. & Eduard, K. TE-greedy-nester: 201. Mousavi, N., Shleizer-Burko, S., Yanicky, R. & Gymrek, M. Profiling the
structure-based detection of LTR retrotransposons and their nesting. genome-wide landscape of tandem repeat expansions. Nucleic Acids Res. 47,
Bioinformatics 36, 4991–4999 (2020). e90 (2019).
172. Wenke, T. et al. Targeted identification of short interspersed nuclear element 202. Dolzhenko, E. et al. ExpansionHunter Denovo: a computational method for
families shows their widespread existence and extreme heterogeneity in plant locating known and novel repeat expansions in short-read sequencing data.
genomes. Plant Cell 23, 3117–3128 (2011). Genome Biol. 21, 1–14 (2020).
173. Hongliang, M. & Hao, W. SINE_scan: an efficient tool to discover short 203. Chiu, R. et al. Straglr: discovering and genotyping tandem repeat expansions
interspersed nuclear elements (SINEs) in large-scale genomic datasets. using whole genome long-read sequences. Genome Biol. 22, 224 (2021).
Bioinformatics 33, 743–745 (2017). 204. Dashnow, H. et al. STRling: a k-mer counting approach that detects short
174. Li, Y., Jiang, N. & Sun, Y. AnnoSINE: a short interspersed nuclear elements tandem repeat expansions at known and novel loci. Genome Biol. 23, 257
annotation tool for plant genomes. Plant Physiol. 188, 955–970 (2022). (2022).
175. Tu, Z. Eight novel families of miniature inverted repeat transposable elements 205. Ou, S. et al. Benchmarking transposable element annotation methods for
in the African malaria mosquito Anopheles gambiae. Proc. Natl. Acad. Sci. creation of a streamlined, comprehensive pipeline. Genome Biol. 20, 275
USA. 98, 1699–1704 (2001). (2019).
206. Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of 229. Rodriguez, M. & Makałowski, W. Software evaluation for de novo detection of
transposable element families. Proc. Natl. Acad. Sci. USA. 117, 9451–9457 transposons. Mobile DNA 13, 1–14 (2022).
(2020). 230. Riehl, K. et al. TransposonUltimate: software for transposon classification,
207. Budiš, J. et al. Dante: genotyping of known complex and expanded short annotation and detection. Nucleic Acids Res. 50, e64–e64 (2022).
tandem repeats. Bioinformatics 35, 1310–1317 (2019). 231. Bell, E. A. et al. Transposable element annotation in non model species: the
208. Abrusán, G., Grundmann, N., DeMester, L. & Makalowski, W. TEclass-a tool benefits of species specific repeat libraries using semi automated EDTA and
for automated classification of unknown eukaryotic transposable elements. DeepTE de novo pipelines. Mol. Ecol. Resour. 22, 823–833 (2022).
Bioinformatics 25, 1329–1330 (2009). 232. Faulk, C. De novo sequencing, diploid assembly, and annotation of the black
209. Hoede, C. et al. PASTEC: an automatic transposable element classification carpenter ant, Camponotus pennsylvanicus, and its symbionts by one person
tool. PLoS ONE 9, e91929 (2014). for $1000, using nanopore sequencing. Nucleic Acids Res. 51, 17–28 (2023).
210. Feschotte, C. et al. Exploring repetitive DNA landscapes using REPCLASS, a 233. Zhang, X., Zhang, R. & Yu, J. New Understanding of the Relevant Role of
tool that automates the classification of transposable elements in eukaryotic LINE-1 Retrotransposition in Human Disease and Immune Modulation.
genomes. Genome Biol. Evol. 1, 205–220 (2009). Front. Cell Dev. Biol. 8, 657 (2020).
211. Mor, B., Garhwal, S. & Kumar, A. A Systematic Review of Hidden Markov
Models and Their Applications. Arch. Computat. Methods Eng. 28, 1429–1448
(2021). Acknowledgements
212. Yan, H., Bombarely, A. & Li, S. DeepTE: a computational method for de novo This work was supported by the King Abdullah University of Science and Technology
classification of transposons with convolutional neural network. (KAUST) Office of Sponsored Research (OSR) under award numbers FCC/1/1976-44-01,
Bioinformatics 36, 4269–4275 (2020). FCC/1/1976-45-01, REI/1/5202-01-01, REI/1/4940-01-01, and RGC/3/4816-01-01, and
213. da Cruz, M. H. P. et al. TERL: classification of transposable elements by the National Natural Science Foundation of China under Grant: No.62002388.
convolutional neural networks. Brief Bioinform. 22, bbaa185 (2021).
214. Martinez-Gomez, L. et al. Few SINEs of life: Alu elements have little evidence Author contributions
for biological relevance despite elevated translation. NAR Genom. Bioinform. X.L., W.Z., and J.Z. researched the literature. X.L., W.Z., J.Z., H.L., X.X., B.Z., and X.G.
2, lqz023 (2020). contributed substantially to discussions of the content. X.L. wrote the paper, and X.G.
215. Salem, A. H. et al. Recently integrated Alu elements and human genomic reviewed and edited the paper.
diversity. Mol. Biol. Evol. 20, 1349–1361 (2003).
216. Hancks, D. C. & Kazazian Jr, H. H. SVA retrotransposons: Evolution and
genetic instability. Semin Cancer Biol. 20, 234–245 (2010). Competing interests
217. Hancks, D. C. et al. The minimal active human SVA retrotransposon requires The authors declare no competing interests.
only the 5’-hexamer and Alu-like domains. Mol. Cell Biol. 32, 4718–4726
(2012).
218. Beck, C. R. et al. LINE-1 retrotransposition activity in human genomes. Cell
Additional information
Supplementary information The online version contains supplementary material
141, 1159–1170 (2010).
available at [Link]
219. Grandi, N. & Tramontano, E. Human Endogenous Retroviruses Are Ancient
Acquired Elements Still Shaping Innate Immune Responses. Front. Immunol.
Correspondence and requests for materials should be addressed to Xin Gao.
9, 2039 (2018).
220. Buzdin, A. et al. Human-specific subfamilies of HERV-K (HML-2) long
Peer review information Communications Biology thanks Indranil Malik and the other,
terminal repeats: three master genes were active simultaneously during
anonymous, reviewer(s) for their contribution to the peer review of this work. Primary
branching of hominoid lineages. Genomics 81, 149–156 (2003).
Handling Editor: George Inglis.
221. van Bree, E. J. et al. A hidden layer of structural variation in transposable
elements reveals potential genetic modifiers in human disease-risk loci.
Reprints and permission information is available at [Link]
Genome Res. 32, 656–670 (2022).
222. Poggi, L. et al. Differential efficacies of Cas nucleases on microsatellites Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in
involved in human disorders and associated off-target mutations. Nucleic published maps and institutional affiliations.
Acids Res. 49, 8120–8134 (2021).
223. Annear, D. J. et al. Non-Mendelian inheritance patterns and extreme deviation
rates of CGG repeats in autism. Genome Res. 32, 1967–1980 (2022).
224. Irigoyen, A. M. et al. Differential expression of the androgen receptor gene is Open Access This article is licensed under a Creative Commons
correlated with CAG polymorphic repeats in patients with prostate cancer. J. Attribution 4.0 International License, which permits use, sharing,
Genet. 102, 23 (2023). adaptation, distribution and reproduction in any medium or format, as long as you give
225. Mu˙ller, N. A. et al. A single gene underlies the dynamic evolution of poplar appropriate credit to the original author(s) and the source, provide a link to the Creative
sex determination. Nat. Plants 6, 630–637 (2020). Commons licence, and indicate if changes were made. The images or other third party
226. Kapitonov, V. V. & Jurka, J. A universal classification of eukaryotic material in this article are included in the article’s Creative Commons licence, unless
transposable elements implemented in Repbase. Nat. Rev. Genet. 9, 411–412 indicated otherwise in a credit line to the material. If material is not included in the
(2008). article’s Creative Commons licence and your intended use is not permitted by statutory
227. Albert, P. S. et al. Whole-chromosome paints in maize reveal rearrangements, regulation or exceeds the permitted use, you will need to obtain permission directly from
nuclear domains, and chromosomal relationships. Proc. Natl. Acad. Sci. USA. the copyright holder. To view a copy of this licence, visit [Link]
116, 1679–1685 (2019). licenses/by/4.0/.
228. Qian, Z. et al. The chromosome level genome of a free floating aquatic weed
Pistia stratiotes provides insights into its rapid invasion. Mol. Ecol. Resour. 22,
2732–2743 (2022). © The Author(s) 2023