Repeatome
Repeatome
45
1. INTRODUCTION
The Human Genome Project initially focused on sequencing the ∼20,000 protein-coding genes,
and there was debate as to whether the rest of the genome, riddled with repetitive sequences,
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53
was worthy of the cost and time to interrogate it. However, we now know that polymorphisms
of interest are often in repeat-rich noncoding regions. Repetitive sequences have been notori-
ously difficult to investigate and map and thus have often been screened out of analysis using
RepeatMasker ([Link] or similar programs. Recent advances in long-
read sequencing technology make it possible to read through and precisely map sequences in
large repetitive regions, allowing comprehensive chromosome sequencing in the Telomere-to-
Telomere (T2T) project (83, 127). This has produced the first full, contiguous sequencing of the
repeatome, including large satellite regions, revealing greater structural complexity than previ-
ously anticipated. Multiple studies have begun to extend this technology to create pangenome
reference sequences that capture polymorphic differences in various repeats in populations (re-
viewed in 155). As we obtain a more complete description of the human genome’s repeat content
and its variations, a major challenge will be to assess the potential impacts of variation in different
types of repeats, and to understand the extent to which various types of repetitive “junk” play any
functional roles.
To undertake a review on the emerging functions of the various abundant repeats in the hu-
man genome is a timely but dauntingly large task. A full accounting of human repeats is beyond
the scope of any review, and we apologize that we cannot fully represent the huge literature of
related work. Here, we present a conceptual overview of emerging functions, focusing largely on
a few of the most abundant repeats with less established functions. Gene regulation has been most
studied in terms of local sequence effects on individual genes, and repeat elements often func-
tion at that level. However, we will convey our perspective that highly abundant repeats may also
function more collectively, to influence regional genome regulation in nuclei. This may best be
understood through the lens of human genome organization on chromosomes and how it relates
to compartmentalized genome regulation within complex nuclear structure.
+ Protein-coding
10 + sequences
(1.50%) SVAs
+
Fraction adsorbed
20 (0.15%)
Low-copy/unique
30 sequences Alu
40 (10.09%)
50
Repetitive Other
DNA Introns SINEs
60 + (26.00%) (2.68%)
70
+ LTRs (8.84%)
80
90 + Other (0.36%)
DNA transposons
100 Housekeeping RNA genes
10–3 10–2 10–1 100 101 102 103 104 (rRNA, tRNA, etc.) (0.11%) Satellites (3.58%)
Simple (4.92%)
C0t (mol × s/L) repeats
(2.54%)
b d LINEs
Non-LTRs
Class I SINEs
retrotransposons
Transposable
elements LTRs HERVs
Interspersed Class II
Density transposons
gradient Satellite band Microsatellites DNA
transposons
Main DNA band Repeats Minisatellites
Macrosatellites
Tandem
Centromeric
satellites
Pericentromeric
satellites
Telomeres
e f
Euchromatin Heterochromatin
Specific gene/
mRNA
Centromeres
1 2 3 4 5 6
Cajal bodies
Nucleolus
7 8 9 10 11 12
rDNA U2
genes
Specific gene/
13 14 15 16 17 18 mRNA
Telomeres Nuclear
Centromeres Nuclear lamina
rDNA speckles
HSat2
HSat3
19 20 21 22 X
(Caption appears on following page)
fraction. Panel adapted with permission from Reference 20. (b) Density gradient centrifugation separating a smaller satellite band from
the main band of genomic DNA. (c) Pie chart showing the relative abundances of repeat types in the human genome. Data are from
Reference 83. (d) Categories of repeat types. Note that not all repeat types in each category are shown here. (e) Karyotype showing
alternating light and dark bands on a G-banded mitotic chromosome spread, with the locations of several repeat sequence types
indicated. Panel adapted with permission from Reference 145. ( f ) Illustration depicting the organization of the interphase nucleus. The
genome has a compartmentalized architecture within the nucleus and is organized further relative to non-membrane-bound
substructures rich in RNA metabolic factors, as indicated. Panel adapted with permission from Reference 145. Abbreviations: HERV,
human endogenous retrovirus; HSat, human satellite; L1, long interspersed nuclear element 1; LINE, long interspersed nuclear
element; LTR, long terminal repeat; SINE, short interspersed nuclear element; SVA, SINE-VNTR-Alu; VNTR, variable number
tandem repeat.
While the potential biological significance for the bulk of repetitive sequences is not known,
numerous studies have shown that a specific repeat sequence, typically near or in a protein-coding
gene, can impact the function of that gene through numerous different mechanisms, which can
be mediated by DNA or RNA. Changes in the location or copy number of a repeat can have
deleterious effects and contribute to disease or can be co-opted during evolution to contribute to
normal gene function. While there are now many examples of a repeat sequence being co-opted to
impact local gene function, they do not necessarily indicate whether the bulk of highly abundant
degenerate repeats contribute to genome function more broadly or if they are just an evolutionary
vestige. We discuss here less established concepts for how certain repeat types, present in enor-
mous numbers (hundreds of thousands to a million), may contribute to the broader regulation of
the genome within nuclear structure.
Identifying and understanding novel mechanisms may require different conceptual approaches
that go beyond the better-known molecular mechanisms that regulate individual genes. Adding
to the challenge, certain repeat types may function only transiently at particular stages of early
development, in response to stress, or in specific disease states, examples of which are mentioned
throughout this review. We consider these functions from the perspective of repeat genome orga-
nization in the human karyotype, which we suggest can provide insight into potentially broader
collective roles of abundant repeats in genome regulation.
pronounced that it is evident from simple staining and light microscopy. In addition to numerous
and often huge pericentric satellites, staining shows a pattern of 400–600 alternating light and
dark Giemsa-stained bands (Figure 1e), which correspond largely to regions with differences in
gene density, short interspersed nuclear elements (SINEs) versus long interspersed nuclear ele-
ments (LINEs), and GC versus AT content. What might be the functional significance of these
cytological-scale differences in linear genome organization? We suggest that a full understanding
will require a perspective on the complex substructure of the interphase nucleus, including the
compartmentalization of euchromatin and heterochromatin into large distinct nuclear regions,
and in cell type–specific patterns (Figure 1f ). The large euchromatin compartment, which is typ-
ically more internal in the nucleus, is punctuated by ∼10–20 discrete nuclear speckles (also known
as SC35 domains) that are concentrated with a host of RNA metabolic factors. This will become
important in Section 4 when we consider the large segmental organization of the genome, as re-
flected in chromosome bands, with differences in the density and types of genes as well as repetitive
sequences.
As alluded to above, the organization of telomere repeats (Figures 1e and 2a) clearly reflects
their function, to cap each chromosome end and protect it from fusing with other chromosomes
(see 28). However, telomere biology also illustrates that a given repeat sequence can have more
than one function and that the study of repeats can reveal unanticipated and fundamentally impor-
tant biology. The discovery that attrition of the telomere array is essentially a cellular aging clock
fueled numerous important discoveries in developmental biology and disease, particularly cancer
(reviewed in 6). Arrays of the telomere repeat TTAGGG are several kilobases in newborns and
are protected by the shelterin complex (reviewed in 43), but telomeres shorten progressively with
each somatic cell division; when they reach a critical length, a DNA damage response then trig-
gers cell senescence. In pluripotent cells, telomere length is maintained by telomerase, an enzyme
largely absent in differentiated cells, leading to telomere shortening and cell senescence (reviewed
in 28, 55). This example affirms the compelling prospects to uncover important new biology by
mining for meaningful information in the complex dark matter of human repetitive sequences.
2.1. Many Mini- and Microsatellite Repeats Can Impact the Functions
of Specific Disease-Associated Genes in Cis
Diverse short tandem repeats (STRs) are present at many loci across our genomes, and changes
in individual tandem repeats can cause dysfunction in specific disease-associated genes. Approxi-
mately 50 monogenic diseases have been linked to the expansion or contraction of STRs (mostly
triplet repeats) in or near disease-causing genes, which can produce gain- or loss-of-function
effects on normal genes. These are primarily neurological diseases such as Huntington disease
(CAG), fragile X syndrome (CGG), myotonic dystrophy (CTG), Friedreich ataxia (GAA),
spinocerebellar ataxia (CAG), and amyotrophic lateral sclerosis (reviewed in 47).
[Link] • Repeat Genome Functions in Nuclear Structure 49
αSat, monomeric/divergent Other
a Telomere b Segmental HSat1 αSat HOR, inactive HSat2 satellite
duplications βSat αSat HOR, active HSat3
Centromere
DNA p arm q arm
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53
Repeat
structure
HSat3B5
D9Z4
Chromosome 9
D7Z2
D7Z1
Large pericentromere
Chromosome 7 Moderate pericentromere
DXZ1
X chromosome Little or no pericentromere
DNA
YTHDC1 HSF1
Normal condition Heat stress Recovery from heat stress
HSat3
Breast tumor
d e
Normal cell Inactive genes Cancer cell
Demethylation
Demethylated
HSat2 at 1q12
sequesters PRC1
in CAP body Loss of ubH2A
Stochastic
expression
of activated HSat2 RNA
1q12 remains genes DNA
transcriptionally
PRC1 and repressed HSat2 RNA PRC1 and
MeCP2 evenly sequesters MeCP2 MeCP2 in
distributed in CAST bodies nuclear bodies
Demethylated
DNA Methylated DNA PRC1
Loss of UbH2A HSat2 RNA MeCP2
(Caption appears on following page)
W. Shay. (b) Schematic of a generalized human peri/centromeric region. The amount of pericentromeric sequence varies greatly among
different chromosomes. Panel adapted with permission from Reference 3. (c) nSB formation. (Top) YTHDC1 (green) and HSF1 (red)
fluorescent immunostaining and DAPI DNA staining (blue) in control cells, recruitment of HSF1 upon heat stress, sequestration of
YTHDC1 at 3 and 6 h during recovery, and return to normal after 24 h. (Bottom) Diagram of HSat3 expression and nSB formation.
HSat3 repeats, which are normally heterochromatin, become expressed upon heat stress. HSat3 RNA recruits specific proteins,
assembling nSBs, which are subsequently remodeled through recruitment of other factors. Top subpanel adapted from Reference 156
(CC BY 4.0); bottom subpanel adapted from Reference 124 (CC BY 4.0). (d) Schematic of CAP and CAST body formation by HSat2.
In many tumors, DNA demethylation triggers HSat2 DNA and RNA molecular sponges, causing further epigenetic dysfunction.
Sequestration of PRC1 at the demethylated 1q12 megasatellite forms CAP bodies and reduces the repressive ubH2A modification at
other HSat2 loci, which express RNA that sequesters MeCP2 in CAST bodies. Panel adapted with permission from Reference 72
(CC BY-NC-ND 4.0). (e) Image showing HSat2 RNA (green) forming CAST bodies in breast tumor cells (with DAPI-stained nuclear
DNA shown in blue). Panel adapted with permission from Reference 72 (CC BY-NC-ND 4.0). Abbreviations: αSat, alpha satellite;
βSat, beta satellite; CAP, cancer-associated Polycomb; CAST, cancer-associated satellite transcript; DAPI, 4′ ,6-diamidino-2-
phenylindole; FISH, fluorescence in situ hybridization; HOR, higher-order repeat; HSat, human satellite; nSB, nuclear stress body;
ubH2A, ubiquitinated histone H2A.
These repeats are also thought to play a variety of roles in normal gene function (reviewed
in 7), including operating as part of gene products [e.g., in coding exons (151) or noncoding
RNA (ncRNA) functional domains (21)], by influencing local chromatin structure and tran-
scription [e.g., nucleosome spacing, CpG methylation, transcription factor (TF) binding sites,
transcription start sites, and enhancers] and acting within untranslated regions or introns to
modulate transcription, translation, and alternative splicing. These highly variable arrays provide a
117). Therefore, STRs could contribute to the missing heritability in multifactorial conditions
and in normal phenotypic variation. And with recent advances in sequencing technology, their
true variation is now being explored (e.g., 137).
We highlight just one of the first triplet repeat disorders discovered, which illustrates a theme
developed further below for major satellites: the capacity of very abundant small repeats to bind
and sequester regulatory factors. Myotonic dystrophy type 1 results from a large expansion of
CTG triplet repeats in the 3′ untranslated region of the DMPK gene. This causes the DMPK
mRNA containing the repeats to accumulate to high levels in the nucleus, forming ribonucleopro-
tein aggregates that sequester an important splicing regulator, MBNL (muscleblind-like). MBNL
levels throughout the nucleoplasm drop sharply as a consequence (144, 167), which impairs alter-
native splicing of pre-mRNAs for many other genes (reviewed in 111). The concept that highly
abundant repeats can bind and impact the distribution of specific nuclear factors will be important
as we consider the function of the much larger tandem repeats, satellites.
(reviewed in 91). The accumulation of ∼100 different proteins on αSat centromeric DNA illus-
trates how high-copy repeats organized into arrays serve to concentrate protein components to
build a structure that functions at that site.
αSat array size and sequence polymorphisms have been associated with defective centromere
architecture and aneuploidies, and polymorphic satellite array size can vary between homologs (re-
viewed in 121, 152). This suggests that these polymorphic differences may play important roles
in human health, but until recently, human centromeres were almost entirely absent from the
genome build, hindering their study. The T2T gapless assembly published in 2022 for the first
time includes all human centromeric sequences, and the inclusion of diverse populations has re-
vealed more variability than expected in αSat sequences, especially among people of recent African
origin (3).
2.4. Diverse Functions for the Huge Pericentric Satellites: HSat2 and HSat3
Pericentric satellites adjacent to the centromere (Figure 2b) were nearly absent from the refer-
ence human genome until recently (2), making them significantly understudied. The two most
abundant are HSat2 and HSat3, which total 28.7 and 47.6 Mb, respectively, and are found on nu-
merous, but not all, human chromosomes. HSat3 is derived from a pentameric repeat, (CATTC)n ,
and the HSat2 repeat is an ∼26-bp degenerate sequence derived from the HSat3 pentamer. To-
gether, they constitute the largest contiguous satellite arrays in the human genome, including an
∼28-Mb HSat3 array on chromosome 9 and the two largest HSat2 arrays, on chromosomes 1 and
16, which are approximately half that size (∼14 Mb) (2) (Figure 1e).
Pericentric satellites are generally silent in most normal cells, with the exception of testis and
brain, and their heterochromatic nature may help stabilize the centromere (reviewed in 60). How-
ever, not all human chromosomes have pericentric satellites (Figure 2b), which indicates that they
are not necessarily required for normal centromere function. Nevertheless, aberrations in their
heterochromatic state or expression have been associated with mitotic defects in spindle attach-
ments and sister chromatid cohesion, as well as increased DNA damage in S-phase due to blocked
replication over pericentric DNA:RNA hybrids (reviewed in 146).
Interestingly, however, pericentric satellites harbor promoter elements that can regulate tran-
scription by RNA polymerase II (RNAPII) or RNAPIII, and pericentric satellite expression is
common during embryogenesis. In fact, many different human satellite families are expressed in
complex patterns during early embryogenesis that appear to be highly regulated (reviewed in 120,
146). This implies directed regulation of individual satellite arrays during specific windows of em-
bryonic development, for currently unknown reasons. One possibility is that pericentric satellite
expression early in embryogenesis plays a role in nucleating the formation of initial heterochro-
matic compartments with this unique chromatin (reviewed in 133). However, this has yet to be
fully explored, as these regions are not currently included in the genome maps for most species
and have only recently been added to the human genome build.
Another possible function is that both HSat2 and HSat3 can impart global gene regulation
through their capacity to act as a cytological-scale molecular sponge. The collective evidence
summarized below suggests that these exceptionally high-copy pericentric arrays, containing
repeat units with protein-binding potential, have extraordinary capacity to amass and cytolog-
ically sequester regulatory factors at both the DNA and RNA levels, thereby modifying their
accessibility on a genome-wide scale. For example, a 14-Mb array of a 26-nt repeat unit will
contain ∼500,000 copies, while a 5-nt sequence could be repeated millions of times in a single
and mouse). This illustrates that the function of repeats is often less stringently tied to primary
sequence than it is for protein-coding genes.
2.4.1. HSat3 DNA and RNA: the nuclear stress sponge. The earliest and most developed
evidence of a human satellite functioning as a sponge is for the very large HSat3 array on chro-
mosome 9 (9q12) and more recently recognized on the Y chromosome. These arrays act at both
the DNA and RNA levels to regulate cell homoeostasis during numerous types of cell stress (e.g.,
heat shock, osmotic or oxidative stress, and UV radiation). Cell stress triggers a series of steps in
which different factors are sequentially bound and released from HSat3 DNA or RNA during the
stress response, including a prolonged recovery process (reviewed in 61, 124) (Figure 2c). At stress
onset, HSF1 (heat shock transcription factor 1) is expressed, and the HSF1 protein localizes to the
normally silent HSat3 loci on chromosome 9 and the Y chromosome, along with several other
TFs and chromatin-remodeling factors. HSF1 activates HSat3 transcription via RNAPII, and the
HSat3 ncRNA transcripts accumulate at the locus, forming ribonucleoprotein bodies called nu-
clear stress bodies. Nuclear stress bodies sequester many different RNA metabolic factors, leading
to global suppression of transcription and translation, until the stress is resolved. During stress re-
covery, these components are released to the nucleoplasm to reactivate the genome in a highly
regulated manner. Nuclear stress bodies remain during the prolonged stress recovery period and
sequester other factors to help reverse the process, suggesting that the same bodies can dynam-
ically change their properties and function. In the final stage of stress recovery, the remaining
HSat3 transcripts recruit repressive factors to re-silence the HSat3 loci, making the sponge inert
once again.
The ability of repetitive RNA to nucleate phased domains confers a second means to affect
genome regulation broadly: by concentrating specific factors together in a reaction crucible that
accelerates biochemical reactions. HSat3 nuclear stress bodies also behave this way during cell
stress (reviewed in 124), illustrating the versatility of satellite RNA bodies in broadly regulating
the genome.
2.4.2. HSat2 DNA and RNA: the nuclear disease sponge. The large HSat2 arrays on chro-
mosomes 1 and 16 are two of the most prominent but poorly studied features of the human
genome. Recent studies have uncovered unanticipated biology of HSat2 satellites that points to
effects mediated by DNA demethylation and RNA expression, but with complex differences be-
tween HSat2 loci on different chromosomes. Translocations and duplications of the large HSat2
array at 1q12 are among the most frequent aberrations in cancers (reviewed in 70), and global DNA
demethylation is also common in human cancers [and in ICF (immunodeficiency, centromeric
region instability, and facial anomalies) syndrome (160)], with HSat2 at 1q12 being especially
sensitive to demethylation (53).
This global demethylation appears to trigger HSat2 at 1q12 to act as a molecular sponge
(72), which further alters the epigenetic state of the cell. DNA demethylation causes PRC1
(Polycomb repressive complex 1), which normally maintains the repressive ubiquitinated his-
tone H2A (UbH2A) mark at target gene loci across the genome, to accumulate over the 1q12
locus, forming large cancer-associated Polycomb (CAP) bodies (72) (Figure 2d). Although sev-
eral human chromosomes have smaller pericentromeric HSat2 arrays, this is a locus-specific
(1q12) and protein-specific (PRC1 but not PRC2) response to DNA demethylation in human
cells, while in mouse, both PRC1 and PRC2 (which trimethylates H3K27) are sequestered to
all major satellites in pericentromeres upon demethylation (33). This difference likely reflects
downstream consequences for global transcriptional regulation, which results in aberrant ex-
pression across the genome (33, 72). Importantly, this includes small HSat2 arrays on other
chromosomes, which become derepressed and aberrantly expressed (72) (Figure 2d). In fact, nu-
merous studies have shown that normally silent HSat2 repeats are commonly overexpressed in
cancers, more frequently than any other satellite (13, 72, 96), as well as during viral infection,
senescence, and DNA damage and in diseases like facioscapulohumeral muscular dystrophy (e.g.,
125, 141).
Aberrant expression from HSat2 loci then compounds the epigenetic dysregulation in these
cells (Figure 2d) by sequestering additional regulatory factors into large ribonucleoprotein bodies.
These HSat2 RNA bodies are prominent hallmarks of many tumors (Figure 2e), detected in ap-
proximately half of 34 diverse tumors examined, and sequester MeCP2 (methyl-CpG binding pro-
tein 2) (72, 102). These were initially termed cancer-associated satellite transcript (CAST) bodies;
however, HSat2 RNA bodies are also seen in other disease contexts, where we call them satellite
transcript (SATT) bodies, and sequester different regulatory factors, including CTCF (CCCTC-
binding factor) (122), EIF4A3 (eukaryotic translation initiation factor 4A3), and ADAR1
(adenosine deaminase RNA 1) (141), leading to further dysregulation of cell homeostasis and gene
expression. Thus, HSat2, like HSat3, can form both DNA and RNA molecular sponges that impact
genome-wide access to important regulatory factors and broadly affect gene expression.
2.4.3. HSat2 expression in development and disease and potential regulatory effects on
gene pathways. Aberrant expression of satellites in disease may not only be a consequence of
misregulation but also directly contribute to effects on specific downstream pathways. There is
some evidence to suggest that cancer cells or viruses co-opt HSat2 expression to impact pathways
that confer a growth advantage. For example, tumors appear to select for expression of specific
satellites (72, 148), and many human viruses use TFs to specifically activate HSat2 loci (126). The
presence of these satellite RNAs (particularly HSat2) is associated with gene expression changes
that confer reduced immune response, changes in cellular motility, or changes in protein stabil-
ity and localization (e.g., 126, 134, 148). And some work has also shown a direct link between
the presence of the HSat2 ncRNA and aberrant expression from specific pathways (loss of the
RNA prevented the effect) (126). This suggests that HSat2 RNA itself altered the regulation of
specific gene pathways via an unknown mechanism, which may well be through sequestration of
their regulatory factors. However, most studies do not look for expression from the repeatome or
sequestration to RNA bodies, so it is unclear whether SATT bodies are responsible for the gene
expression changes observed.
Most human satellite families are enriched in a wide variety of satellite-specific TF binding
sites, including those that regulate conserved signaling pathways (62, 162). Since expression of
different satellite families appears to be highly regulated throughout embryogenesis, and global
hypomethylation of satellites is also a normal hallmark of gametes, preimplantation embryos, and
extraembryonic tissues (reviewed in 163), it is tempting to speculate that satellites may act as DNA
or RNA molecular sponges at important transition points during normal development. This may
dynamically regulate global genomic access to specific regulatory factors during developmental
transitions.
The likelihood that satellites can act as DNA or RNA sponges during embryogenesis is
supported by findings that mouse pericentromeres (which form chromocenters) can functionally
sequester specific TFs during mouse embryogenesis (110). Additionally, the DUX4 (double ho-
meobox 4) TF is expressed during specific windows of human embryogenesis [and by many human
regulators.
The fascinating story of DUX4 began with the discovery of its abnormal activation in fa-
cioscapulohumeral muscular dystrophy (50) and illustrates the importance of including the
repeatome in both transcriptomic and molecular cytology studies (reviewed in 123). The
macrosatellite D4Z4, on the subtelomere of chromosome 4q, is heterochromatic and silent in
most adult cell types. However, reducing the copy number of the 3.3-kb repeat in the D4Z4 array
to <10 triggers loss of heterochromatic repression and aberrant DUX4 expression, which is con-
sidered the primary cause of muscle degeneration (71). DUX4, which normally regulates HSat2
expression during development, induces aberrant HSat2 RNA in facioscapulohumeral muscular
dystrophy muscle. HSat2 SATT bodies in cell nuclei sequester important regulatory factors that
affect RNA stability, splicing, and translation. Hence, future studies will need to consider what
downstream consequences are directly due to the DUX4 TF or, alternatively, might be due to
effects of sequestration of regulatory factors by HSat2.
This example illustrates how a macrosatellite can act locally to regulate the epigenetic state of a
locus (4q35) that encodes an important regulatory gene (DUX4), which in turn normally regulates
embryonic expression of a major satellite (HSat2), potentially affecting nuclear regulatory factors
more globally.
Poly(A)
L1 5’ UTR ORF1 EN
ORF2
RT C 3’ UTR tail
ORF0
Active TEs TEDS
~300 nt
Left arm Right arm New regulatory
Deletions elements
Alu A B AAA AAAAAA
Co-opted
Duplications New proteins for gene
function
b Active TEs
(0.02%)
Rearrangements New lncRNAs Genome
Disease evolution
Untreated
nuclei
Extracted
nuclei
f g
Percentage of the genome
Genome DNA
15 Scaffold RNA
XIST RNA hCOT1 RNA
Soluble RNA
DAPI DNA DAPI DNA
10
0
HT1080 G3 GM11687 hybrid
L1
L2
IR
V
SR
ou r/
hA
Al
ER
gu e
M
s
/
bi th
LC
O
am
h
L1 DNA Alu DNA CpG DNA Late-replicating DNA
showing the typically negative effects of TE transposition compared with the positive contributions of TEDS in the function of
individual genes, as well as their emerging broader role in nuclear genome architecture. (d) Line drawings of plasma cell and monocyte
nuclei illustrating how the organization of the condensed heterochromatic compartment and the more open euchromatin differs
between cell types. Panel adapted with permission from Reference 26. (e) Images of untreated and extracted nuclei showing that both
XIST RNA (red) and C0 t-1 RNA (green) remain localized and bound with the nuclear scaffold after nuclear extraction and removal of
histones and DNA. Panel adapted with permission from Reference 39. (f ) XIST RNA in HT1080 G3 cells (left) and C0 t-1 RNA in
GM11687 hybrid cells (right). Similar to how XIST RNA localizes to the inactive X chromosome territory (left), human C0 t-1 RNA
strictly localizes on the active human chromosome territory in hybrid cells (right). Panel adapted with permission from Reference 73.
(g) Graph of soluble or scaffold-associated repeat RNAs relative to their abundance in the genome. Repeats sequenced in nuclear RNA
are overwhelmingly associated with the insoluble nuclear scaffold. Panel adapted with permission from Reference 39. (h, left) Examples
of L1 and Alu distribution on several human mitotic chromosomes as detected by DNA FISH. (Right) The same chromosomes labeled
for CpG density and late replication. Left subpanel adapted with permission from Reference 74; right subpanel adapted with
permission from Reference 15. Abbreviations: C, cysteine-rich domain; DAPI, 4′ ,6-diamidino-2-phenylindole; EN, endonuclease; ERV,
endogenous retrovirus; FISH, fluorescence in situ hybridization; L1/2, long interspersed nuclear element 1/2; LC, low complexity;
lncRNA, long noncoding RNA; MIR, mammalian-wide interspersed repeat; ORF, open reading frame; RT, reverse transcriptase; SR,
simple repeats; TE, transposable element; TEDS, TE-derived sequences; UTR, untranslated region.
humans, by altering the structure or regulation of specific genes and triggering chromosomal
rearrangements (58).
3.2. Transposition Activity of Intact LINEs and SINEs and Its Consequences
Less than 0.05% of the millions of TE sequences remain intact and capable of transposition
(Figure 3b), and all of them are retrotransposons, including evolutionarily young families of intact
LINEs and SINEs. The rate of transposition in humans is low: Approximately 1 in every 17 births
carries a new TE integration (59). Since TE activation can have harmful effects, mobile TEs are
largely silenced via epigenetic mechanisms, including DNA methylation, histone modifications,
and silencing mediated by Piwi-interacting RNA and small interfering RNA (1, 34). While tightly
regulated in most somatic tissues, L1 expression and retrotransposition do occur at specific times
in development (154). Transposition of L1 is more frequent in early embryogenesis and occurs in
specialized cells such as spermatozoa and oocytes (65, 105). Low-level L1 and/or Alu activity may
also contribute to somatic mosaicism (92), particularly in the brain (reviewed in 16).
gave rise to millions of TEDS, which for poorly understood reasons have persisted to vastly
outnumber the active TEs (Figure 3b). Some TEDS have been domesticated for normal gene
functions, such as generating new functional regulatory elements, ncRNAs, and proteins (reviewed
in 56, 63) (Figure 3c). For instance, more than 20% of regulatory elements in the human genome
are TE derived, and more than 85% of these are primate specific (4, 51). Approximately 75%
of human genes have at least one Alu sequence, and examples of Alu regulating the function of
nearby genes are especially numerous. These include acting as cis-acting DNA regulatory ele-
ments (e.g., promoters, enhancers, insulators, or TF binding sites) or within mRNAs (in introns
or untranslated regions) to influence splicing, nuclear retention, and mRNA stability (reviewed in
170). A recent study showed that some enhancers may use RNA pairing to interact with specific
promoters and that almost 40% of these RNA interaction sites overlap Alu sequences (108). In ad-
dition, transcription of L1 sequences has enhancer functions that are essential to zygotic genome
activation in mouse embryos (107).
TE sequences are also a source for the evolution of new genes (Figure 3c). More than
80% of human long ncRNAs (lncRNAs) contain at least one TEDS, with TEDS compris-
ing ∼40% of lncRNA sequences (93). For example, the structural RNAs NEAT1 and XIST,
which are responsible for the formation of paraspeckles and X inactivation/Barr body forma-
tion, respectively, contain numerous repetitive sequences, some of which may derive from TEs,
and which serve as binding sites for proteins essential to their function (54, 169). In addi-
tion to lncRNAs, microRNAs and Piwi-interacting RNAs can also be derived from TEs. TEs
have also been exapted to create more than 100 new proteins. For instance, CENP-B, which
is involved in centromere formation, was derived from a DNA transposon. A variety of pro-
teins important in lymphocyte, placenta, and brain development are TE derived (reviewed in
56).
The studies cited above and numerous others have demonstrated that an interspersed repeat
sequence can contribute to the regulation or function of a nearby gene, or as part of the gene
itself. However, this does not necessarily attribute functionality to the sea of innumerable re-
peats interlaced through the whole genome. A major challenge remains to understand whether the
abundance of interspersed repeats serves some general genomic function or is mostly evolutionary
detritus. This question is the focus of the next section.
and 3d): condensed, inactive heterochromatin, mostly near the nuclear or nucleolar peripheries,
and open euchromatin, which occupies much of the interior nuclear regions in most cell types.
It may often be thought that the activity of individual genes explains the visible decondensation
evident throughout euchromatin, but it does not. For perspective, it is important to recognize
issues of scale. Some (but not all) genes within the decondensed euchromatic compartment will
be expressed, but this is a small fraction of the total open chromatin in this region, and pack-
aging changes for individual active genes occur at a much smaller scale than the formation of
the compartment. Increasing evidence supports that regional formation of heterochromatin is
not driven by the off state of individual genes. For example, during initiation of X inactivation
(induced by XIST RNA), the large, condensed Barr body forms before chromosome-wide gene
silencing (157), and Polycomb complexes (PRC1) can mediate long-range DNA interactions to
form heterochromatin compartments independent of local histone modifications and gene repres-
sion (18). Similarly, the initiation of the nuclear heterochromatic compartment forms in two- to
four-cell embryos before any cell type–specific changes in gene expression (25, 85).
These cytologically visible nuclear compartments are at a larger scale than structures detected
by Hi-C (chromosome conformation capture), which uses DNA cross-linking to investigate se-
quence organization in nuclei. This approach identifies topologically associating domains (TADs)
or sub-TADs and the larger A and B compartments (109). TADs are small intrachromosomal
self-interacting regions that are tethered by CTCF binding sites. Notably, more than 95% of
mammalian CTCF sites are derived from TEs (SINEs, LINEs, and long terminal repeats), and
almost all disease-associated STRs localize with CTCF boundaries (79). Although recent studies
have found increasing complexity to Hi-C structures, they are not clearly linked to euchro-
matin/heterochromatin packaging. However, TADs are bundled into A and B compartments
(of variable size, ∼1 Mb or more), which correspond to euchromatin (A) and heterochromatin
(B1/B2) bundles. While each A and B compartment reflects packaging well above the gene level,
the cytological-scale nuclear compartments are built by congregation of numerous A-with-A and
B-with-B Hi-C compartments.
We hypothesize below that the repetitive sequences that make up much of the fabric of a
chromosomal region are related to, and likely play a role in, forming heterochromatin versus
euchromatin regions. Before discussing how this relates to the karyotypic organization of gene
and repeat sequences, we summarize the important point that repeats are abundant not only in
DNA but also in nuclear RNA.
C0 t-1 RNA territory remained after prolonged transcriptional inhibition but could be rapidly dis-
persed by disrupting a nuclear scaffold protein, causing chromatin condensation (e.g., 98), which
suggested that it could be a euchromatic structural RNA. This idea was supported by several
studies showing that disruption of nuclear RNA causes cytological chromatin condensation and
implicating HnRNP-U (heterogeneous nuclear ribonucleoprotein U)/SAF-A (scaffold attach-
ment factor A) or similar proteins, which have both DNA- and RNA-binding domains, as being
involved (reviewed in 115).
To identify the RNA sequences involved in nuclear chromosome structure, a biochemical frac-
tionation procedure was developed to isolate nuclear scaffold RNAs that remain insoluble after
removal of histones and DNA (39) (Figure 3e). The procedure extracts most nuclear RNA and
leaves just 15% that cofractionates with known architectural RNAs, XIST RNA, and NEAT1
RNA [which forms the scaffold for nuclear paraspeckles (32)]. The insoluble RNAs that remained
with the nuclear scaffold are composed almost entirely of long, repeat-rich C0 t-1 RNAs (pre-
mRNAs, lncRNAs, and lincRNAs), and repeat RNA sequences are found almost entirely in the
nuclear scaffold fraction (Figure 3g). This C0 t-1 heterogeneous nuclear RNA is associated with
known nuclear scaffold/matrix RNA-binding proteins [matrin 3, NuMa (nuclear mitotic appara-
tus), and SAF-A] and forms an RNA-binding protein meshwork that promotes open chromatin.
A recent report also directly showed that matrin 3 binds repeat RNAs, particularly L1 anti-
sense RNA, and that the disruption of RNA binding (by a mutation that causes amyotrophic
lateral sclerosis) causes aberrant chromatin condensation (171). Thus, evidence supports that long,
repeat-rich “junk” RNA is integral to maintaining euchromatin structure and that the repeats
within this RNA may play a key role.
This evidence that long, repeat-rich heterogeneous nuclear RNAs function in maintaining
open chromatin suggests an unanticipated role for intron sequences in euchromatin structure
around active genes. Introns often allow for alternative splicing; however, this does not explain
their excessive length [some over 50 kb (142)] or why 80–90% of pre-mRNA sequence (and
many lncRNAs) is noncoding and replete with repeats. Repeat-rich intronic RNAs, lncRNAs, and
lincRNAs might help stabilize the epigenetic state of euchromatin and may also explain recent
findings that revealed exceptionally long-lived RNAs, including pre-mRNAs and lncRNAs, in nu-
clei of terminally differentiated mouse neurons (173); these RNAs appear to be highly analogous
to human C0 t-1 scaffold RNAs (for commentary, see 104).
Interspersed repeat RNAs may also play protective roles throughout the nuclear genome in
response to stress, adding to the evidence of a function for repeats in the stress response, as
established for HSat3 RNA (detailed in Section 2.4.1). For example, in response to stress, Alu ele-
ments are expressed from their own RNAPIII promoter, and the transcripts directly bind RNAPII,
repressing global transcription (116). The subsequent widespread loss of transcription would oth-
erwise cause deleterious chromatin condensation, but, surprisingly, new transcription of long,
intergenic, repeat-rich C0 t-1 RNAs is induced upon stress, including in response to osmotic shock
(159) or reversible transcriptional arrest (39). These extremely long intergenic transcripts have
been termed DOGS (downstream of genes) and suggested to play a role in protecting chromatin
from collapse by maintaining euchromatic C0 t-1 RNAs. This idea was supported by the observa-
tion that despite the arrest of genic transcription during the stress response, the new intergenic
transcription maintained C0 t-1 scaffold RNA levels (39), and this was required to avoid chromatin
collapse. These findings highlight that different components of the noncoding repeatome play a
role in response to stress.
Abbreviations: LINE, long interspersed nuclear element; SINE, short interspersed nuclear element; TE, transposable
element.
are expressed in close proximity to nuclear speckles (143). Gene clustering around nuclear hubs
that promote efficient gene expression further provides a functional rationale for the large-scale
clustered regional distribution of coding genes on chromosomes. This demonstrates a fundamen-
tal relationship between the segmental organization of the linear genome on chromosomes and
the structural organization of the functional genome in nuclei. It also explains why Alu-rich DNA
is more densely clustered around these same structures (29, 74) and why Alu-rich DNA is highly
correlated with regions of highest expression in the interphase nucleus (29).
Since Alu SINEs can contribute to the functions of individual genes, their enrichment in
gene-rich R-bands may simply reflect evolutionary conservation. However, their presence in gene-
rich regions would not require that Alu be strongly depleted from other regions (discussed in
Section 4.4), and the 1.1 million Alu TEDS that are more concentrated in the gene-rich segments
suggest greater Alu density than is easily explained by individual gene regulation. Hence, there
remains a question as to whether regional densities of Alu may be evolutionarily conserved for
additional, perhaps broader, contributions.
SC35
L1-rich
chromatin
b A B
c
100 Tig-1
Percentage of repeat-rich
18%
80
compartments
60 77%
82%
40
20
23%
0
B1 rich L1 rich
L1 DNA Alu DNA L1 DNA Alu DNA
log2(B1/L1) >0 <0
Compartments called de novo
(p < 4 × 10−15)
d e 30
L1 DNA Alu DNA X chromosome
L1 LINEs (percentage of sequence)
Highest L1
25
Chromosome 13
15
interactions
Long-range
Chromosome 20
0.8
Chromosome 21 Chromosome 17
0.4 Y chromosome Chromosome 16
10
0.0 Chromosome 19
Chromosome 22 Highest Alu
60 5
content (%) content (%)
Alu rich 0 5 10 15 20 25
40 L1 rich Alu SINEs (percentage of sequence)
Alu
20
0
60
40
L1
20
0
20 30 40
Chromosomal location (Mb)
(Caption appears on following page)
periphery (SC35 speckles; green). (Right) Model showing that gene- and Alu-rich R-band DNA (light blue) is more intimately associated
with nuclear speckles than L1-rich, gene-poor G-band DNA (dark blue). Panel adapted with permission from Reference 143.
(b) Percentages of repeat-rich compartments as examined by Hi-C approaches. The most SINE-rich (B1 in mouse) compartments are
primarily A compartments (active), whereas the most LINE-rich (L1) are B compartments (inactive). Panel adapted from Reference
112 (CC BY 4.0). (c) DNA FISH images for L1 (red) and Alu (green) in human fibroblast nucleus, with the green signal outlined in the
second image to show Alu depletion at the periphery (DAPI; blue). Panel adapted with permission from Reference 74. (d, top) DNA
FISH images of a senescent fibroblast nucleus for L1 (red) and Alu (green) DNA. On the right, DAPI DNA shows dense
heterochromatin foci (SAHFs). (Bottom) Ideogram of 50 Mb of chromosome 4 aligned with the graph, showing changes in long-range
(>10 Mb) intrachromosomal Hi-C interactions between senescence and growing cells (log2 ). Each dot represents 100 kb and is colored
according to Giemsa-band designations; the black arrow indicates a condensing L1-rich R-band, and the blue arrow indicates an
R-band containing Alu-rich peaks that resist condensation. Also shown are the percentages of Alu and L1 content across the same
region, with the highest (90th percentile) Alu content outlined in green and the highest L1 content outlined in blue. Panel adapted with
permission from Reference 74. (e) Relative contributions (as percentages of total chromosome sequence) for L1 and Alu in all human
chromosomes. The insets show the X chromosome and chromosome 19 stained for L1 (red) and Alu (green) DNA. Panel adapted with
permission from Reference 74. Abbreviations: DAPI, 4′ ,6-diamidino-2-phenylindole; FISH, fluorescence in situ hybridization; L1, long
interspersed nuclear element 1; LINE, long interspersed nuclear element; SAHF, senescence-associated heterochromatic focus; SINE,
short interspersed nuclear element.
Recent evidence indicates that the depletion of Alu-rich peaks, not just L1 density, may be
an important factor influencing whether there is condensation of a region (74). In many primary
human fibroblasts, the peripheral heterochromatin appears to be more clearly delineated by Alu
depletion than by L1 enrichment (74) (Figure 4c). The condensed L1-rich Barr body, at the core of
the inactive X chromosome, also excludes Alu-rich DNA. In addition, when senescent cells com-
pletely reorganize peripheral heterochromatin into senescence-associated heterochromatic foci
(SAHFs), Alu-rich regions are excluded from these condensed bodies as well (74) (Figure 4d).
Hi-C data analysis revealed long contiguous Alu peaks as the most striking variance in repeat dis-
tribution, and the Alu-peak region consistently countered chromosome compaction (as indicated
by increased long-range intrachromosomal interactions). Figure 4d shows a quantification of L1
and Alu (in 100-kb bins) that demonstrates two additional points. First, L1 density on this chro-
mosome (chromosome 4) is similarly high across both R- and G-bands, which can show similar
structural changes (e.g., the black arrow indicates a condensing L1-rich R-band); in contrast, the
R-band containing Alu-rich peaks (blue arrow) resists this change (condensation). Second, results
also show that architectural interactions change in unison across whole large (∼5–15 Mb) chro-
mosome segments; for example, DNA throughout the whole darkest G-band shows increased
long-range interactions, suggesting that the band’s architecture changes as a single structural
unit.
Evidence indicates that constitutive heterochromatin forms in the darkest G-bands (G-positive
bands 75–100), which are the most L1 rich but also the lowest in Alu. Data from an earlier study of
five different band classes (64) indicated that SINEs (as a percentage of total sequence) essentially
double between the darkest and lightest bands (from 8.4% to 15.6%), whereas LINEs decrease
by ∼25% (from 25.1% to 19.2%). Marked differences in Alu and L1 enrichment are also seen
for certain whole chromosomes that similarly differ in their propensity to form heterochromatin.
The X chromosome has the highest L1 content (Figure 4e), although the L1 density is lower
in the pseudoautosomal region that escapes gene silencing (8). In marked contrast, chromosome
19 is strikingly Alu rich and low in L1, has the highest gene density, and consistently resides in
the euchromatic nuclear interior (78). Chromosome 19 is also unusual in that it does not form
a heterochromatic SAHF in senescent cells (74). Furthermore, chromosome 19 is an outlier in
that it encodes a concentration of more than 250 zinc finger regulatory proteins (44), many of
DISCLOSURE STATEMENT
The authors are not aware of any affiliations, memberships, funding, or financial holdings that
might be perceived as affecting the objectivity of this review.
ACKNOWLEDGMENTS
We appreciate the support of National Institutes of Health grant R35 GM122597 to J.B.L.
LITERATURE CITED
1. Almeida MV, Vernaz G, Putman ALK, Miska EA. 2022. Taming transposable elements in vertebrates:
from epigenetic silencing to domestication. Trends Genet. 38:529–53
2. Altemose N. 2022. A classical revival: human satellite DNAs enter the genomics era. Semin. Cell Dev.
Biol. 128:2–14
3. Altemose N, Logsdon GA, Bzikadze AV, Sidhwani P, Langley SA, et al. 2022. Complete genomic and
epigenetic maps of human centromeres. Science 376:eabl4178
4. Andrews G, Fan K, Pratt HE, Phalke N, Zoonomia Consort., et al. 2023. Mammalian evolution of human
cis-regulatory elements and transcription factor binding sites. Science 380:eabn7930
5. Ardeljan D, Steranka JP, Liu C, Li Z, Taylor MS, et al. 2020. Cell fitness screens reveal a conflict between
LINE-1 retrotransposition and DNA replication. Nat. Struct. Mol. Biol. 27:168–78
6. Armanios M. 2022. The role of telomeres in human disease. Annu. Rev. Genom. Hum. Genet. 23:363–81
7. Bagshaw ATM. 2017. Functional mechanisms of microsatellite DNA in eukaryotic genomes. Genome
Biol. Evol. 9:2428–43
8. Bailey JA, Carrel L, Chakravarti A, Eichler EE. 2000. Molecular evidence for a relationship between
LINE-1 elements and X chromosome inactivation: the Lyon repeat hypothesis. PNAS 97:6634–39
9. Bak AL, Jorgensen AL, Zeuthen J. 1981. Chromosome banding and compaction. Hum. Genet. 57:199–
202
10. Batzer MA, Deininger PL. 2002. Alu repeats and human genomic diversity. Nat. Rev. Genet. 3:370–79
11. Beck CR, Garcia-Perez JL, Badge RM, Moran JV. 2011. LINE-1 elements in structural variation and
disease. Annu. Rev. Genom. Hum. Genet. 12:187–215
12. Belancio VP, Hedges DJ, Deininger P. 2006. LINE-1 RNA splicing and influences on mammalian gene
expression. Nucleic Acids Res. 34:1512–21
13. Bersani F, Lee E, Kharchenko PV, Xu AW, Liu M, et al. 2015. Pericentromeric satellite repeat expansions
through RNA-derived DNA intermediates in cancer. PNAS 112:15148–53
14. Betancourt AJ, Wei KH, Huang Y, Lee YCG. 2024. Causes and consequences of varying transposable
element activity: an evolutionary perspective. Annu. Rev. Genom. Hum. Genet. 25:1–25
15. Bickmore WA. 2019. Patterns in the genome. Heredity 123:50–57
16. Bizzotto S. 2023. The human brain through the lens of somatic mosaicism. Front. Neurosci. 17:1172469
17. Bourque G, Burns KH, Gehring M, Gorbunova V, Seluanov A, et al. 2018. Ten things you should know
about transposable elements. Genome Biol. 19:199
18. Boyle S, Flyamer IM, Williamson I, Sengupta D, Bickmore WA, Illingworth RS. 2020. A central role
for canonical PRC1 in shaping the 3D nuclear landscape. Genes Dev. 34:931–49
19. Britten RJ, Davidson EH. 1969. Gene regulation for higher cells: a theory. Science 165:349–57
20. Britten RJ, Kohne DE. 1968. Repeated sequences in DNA: Hundreds of thousands of copies of DNA
sequences have been incorporated into the genomes of higher organisms. Science 161:529–40
21. Brockdorff N. 2018. Local tandem repeat expansion in Xist RNA as a model for the functionalisation of
ncRNA. Noncoding RNA 4:28
100. Korenberg JR, Rykowski MC. 1988. Human genome organization: Alu, LINES, and the molecular
structure of metaphase chromosome bands. Cell 53:391–400
101. Lander ES, Linton LM, Birren B, Nusbaum C, Zody MC, et al. 2001. Initial sequencing and analysis of
the human genome. Nature 409:860–921
102. Landers CC, Rabeler CA, Ferrari EK, D’Alessandro LR, Kang DD, et al. 2021. Ectopic expression of
pericentric HSATII RNA results in nuclear RNA accumulation, MeCP2 recruitment, and cell division
defects. Chromosoma 130:75–90
103. Larsen PA, Hunnicutt KE, Larsen RJ, Yoder AD, Saunders AM. 2018. Warning SINEs: Alu elements,
evolution of the human brain, and the spectrum of neurological disease. Chromosome Res. 26:93–111
104. Lawrence J, Hall L. 2024. Exceptionally long-lived nuclear RNAs. Science 384:31–32
105. Lazaros L, Kitsou C, Kostoulas C, Bellou S, Hatzi E, et al. 2017. Retrotransposon expression
and incorporation of cloned human and mouse retroelements in human spermatozoa. Fertil. Steril.
107:821–30
106. Li S, Shen X. 2023. Long interspersed nuclear element 1 and B1/Alu repeats blueprint genome
compartmentalization. Curr. Opin. Genet. Dev. 80:102049
107. Li X, Bie L, Wang Y, Hong Y, Zhou Z, et al. 2024. LINE-1 transcription activates long-range gene
expression. Nat. Genet. 56:1494–502
108. Liang L, Cao C, Ji L, Cai Z, Wang D, et al. 2023. Complementary Alu sequences mediate enhancer-
promoter selectivity. Nature 619:868–75
109. Lieberman-Aiden E, van Berkum NL, Williams L, Imakaev M, Ragoczy T, et al. 2009. Comprehensive
mapping of long-range interactions reveals folding principles of the human genome. Science 326:289–93
110. Liu X, Wu B, Szary J, Kofoed EM, Schaufele F. 2007. Functional sequestration of transcription factor
activity by repetitive DNA. J. Biol. Chem. 282:20868–76
111. López-Martínez A, Soblechero-Martín P, de-la-Puente-Ovejero L, Nogales-Gadea G, Arechavala-
Gomeza V. 2020. An overview of alternative splicing defects implicated in myotonic dystrophy type I.
Genes 11:1109
112. Lu JY, Chang L, Li T, Wang T, Yin Y, et al. 2021. Homotypic clustering of L1 and B1/Alu repeats
compartmentalizes the 3D genome. Cell Res 31:613–30
113. Lu JY, Shao W, Chang L, Yin Y, Li T, et al. 2020. Genomic repeats categorize genes with distinct
functions for orchestrated regulation. Cell Rep. 30:3296–311.e5
114. Lyon MF. 2000. LINE-1 elements and X chromosome inactivation: a function for “junk” DNA? PNAS
97:6248–49
115. Marenda M, Lazarova E, Gilbert N. 2022. The role of SAF-A/hnRNP U in regulating chromatin
structure. Curr. Opin. Genet. Dev. 72:38–44
116. Mariner PD, Walters RD, Espinoza CA, Drullinger LF, Wagner SD, et al. 2008. Human Alu RNA is a
modular transacting repressor of mRNA transcription during heat shock. Mol. Cell 29:499–509
117. Marshall JN, Lopez AI, Pfaff AL, Koks S, Quinn JP, Bubb VJ. 2021. Variable number tandem repeats—
their emerging role in sickness and health. Exp. Biol. Med. 246:1368–76
118. McKerrow W, Wang X, Mendez-Dorantes C, Mita P, Cao S, et al. 2022. LINE-1 expression in cancer
correlates with p53 mutation, copy number alteration, and S phase checkpoint. PNAS 119:e2115999119
119. McNeil JA, Smith KP, Hall LL, Lawrence JB. 2006. Word frequency analysis reveals enrichment of
dinucleotide repeats on the human X chromosome and [GATA]n in the X escape region. Genome Res.
16:477–84
120. Miga KH. 2019. Centromeric satellite DNAs: hidden sequence variation in the human population. Genes
10:352
121. Miga KH, Alexandrov IA. 2021. Variation and evolution of human centromeres: a field guide and
perspective. Annu. Rev. Genet. 55:583–602
122. Miyata K, Imai Y, Hori S, Nishio M, Loo TM, et al. 2021. Pericentromeric noncoding RNA
changes DNA binding of CTCF and inflammatory gene expression in senescence and cancer. PNAS
118:e2025647118
repeat element polarization between immunotherapy responsive and T cell suppressive classes. Cell Rep.
23:512–21
149. Stamidis N, Zylicz JJ. 2023. RNA-mediated heterochromatin formation at repetitive elements in
mammals. EMBO J. 42:e111717
150. Stein RA, DePaola RV. 2023. Human endogenous retroviruses: our genomic fossils and companions.
Physiol. Genom. 55:249–58
151. Stoyas CA, La Spada AR. 2018. The CAG-polyglutamine repeat diseases: a clinical, molecular, genetic,
and pathophysiologic nosology. Handb. Clin. Neurol. 147:143–70
152. Sullivan LL, Sullivan BA. 2020. Genomic and functional variation of human centromeres. Exp. Cell Res.
389:111896
153. Sultana T, van Essen D, Siol O, Bailly-Bechet M, Philippe C, et al. 2019. The landscape of L1 retrotrans-
posons in the human genome is shaped by pre-insertion sequence biases and post-insertion selection.
Mol. Cell 74:555–70.e7
154. Talley MJ, Longworth MS. 2024. Retrotransposons in embryogenesis and neurodevelopment. Biochem.
Soc. Trans. 52:1159–71
155. Taylor DJ, Eizenga JM, Li Q, Das A, Jenike KM, et al. 2024. Beyond the Human Genome Project: the
age of complete human genome sequences and pangenome references. Annu. Rev. Genom. Hum. Genet.
25:77–104
156. Timcheva K, Dufour S, Touat-Todeschini L, Burnard C, Carpentier MC, et al. 2022. Chromatin-
associated YTHDC1 coordinates heat-induced reprogramming of gene expression. Cell Rep. 41:111784
157. Valledor M, Byron M, Dumas B, Carone DM, Hall LL, Lawrence JB. 2023. Early chromosome
condensation by XIST builds A-repeat RNA density that facilitates gene silencing. Cell Rep. 42:112686
158. Vergnaud G, Denoeud F. 2000. Minisatellites: mutability and genome architecture. Genome Res. 10:899–
907
159. Vilborg A, Steitz JA. 2017. Readthrough transcription: How are DoGs made and what do they do? RNA
Biol. 14:632–36
160. Vukic M, Daxinger L. 2019. DNA methylation in disease: immunodeficiency, centromeric instability,
facial anomalies syndrome. Essays Biochem. 63:773–83
161. Vuoristo S, Bhagat S, Hyden-Granskog C, Yoshihara M, Gawriyski L, et al. 2022. DUX4 is a
multifunctional factor priming human embryonic genome activation. iScience 25:104137
162. Wada Y, Iwasaki Y, Abe T, Wada K, Tooyama I, Ikemura T. 2015. CG-containing oligonucleotides
and transcription factor-binding motifs are enriched in human pericentric regions. Genes Genet. Syst.
90:43–53
163. Walton EL, Francastel C, Velasco G. 2014. Dnmt3b prefers germ line genes and centromeric regions:
lessons from the ICF syndrome and cancer and implications for diseases. Biology 3:578–605
164. Wang J, Vicente-Garcia C, Seruggia D, Molto E, Fernandez-Minan A, et al. 2015. MIR retrotransposon
sequences provide insulators to the human genome. PNAS 112:E4428–37
165. Waring M, Britten RJ. 1966. Nucleotide sequence repetition: a rapidly reassociating fraction of mouse
DNA. Science 154:791–94
166. Wells JN, Feschotte C. 2020. A field guide to eukaryotic transposable elements. Annu. Rev. Genet. 54:539–
61
167. Wheeler TM, Thornton CA. 2007. Myotonic dystrophy: RNA-mediated muscle disease. Curr. Opin.
Neurol. 20:572–76
168. Xie Z, Liu C, Lu Y, Sun C, Liu Y, et al. 2022. Exonization of a deep intronic long interspersed nuclear
element in Becker muscular dystrophy. Front. Genet. 13:979732
169. Yamazaki T, Souquere S, Chujo T, Kobelke S, Chong YS, et al. 2018. Functional domains of NEAT1
architectural lncRNA induce paraspeckle assembly through phase separation. Mol. Cell 70:1038–53.e7
170. Zhang XO, Pratt H, Weng Z. 2021. Investigating the potential roles of SINEs in the human genome.
Annu. Rev. Genom. Hum. Genet. 22:199–218