0% found this document useful (0 votes)
7 views31 pages

Repeatome

The document discusses the emerging functions of repetitive sequences in the human genome, which make up about half of its composition. It highlights how these sequences, often considered 'junk' DNA, may play significant roles in nuclear structure and genome regulation, particularly through their organization within the human karyotype. Advances in sequencing technology are enabling a deeper understanding of these repeat sequences and their potential collective functions beyond individual gene regulation.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views31 pages

Repeatome

The document discusses the emerging functions of repetitive sequences in the human genome, which make up about half of its composition. It highlights how these sequences, often considered 'junk' DNA, may play significant roles in nuclear structure and genome regulation, particularly through their organization within the human karyotype. Advances in sequencing technology are enabling a deeper understanding of these repeat sequences and their potential collective functions beyond individual gene regulation.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: 14.139.123.

36 On: Sun, 11 Jan 2026 15:19:53

Annual Review of Genomics and Human Genetics


Emerging Functions of the
Repeat Genome in Nuclear
Structure: A View from the
Human Karyotype
Lisa L. Hall, Kelly P. Smith, and Jeanne B. Lawrence
Department of Neurology, University of Massachusetts Chan Medical School, Worcester,
Massachusetts, USA; email: [Link]@[Link]

Annu. Rev. Genom. Hum. Genet. 2025. 26:45–75 Keywords


First published as a Review in Advance on
human repeats, satellites, transposable elements, noncoding RNA,
May 29, 2025
chromosome and nuclear structure, heterochromatin, euchromatin
The Annual Review of Genomics and Human Genetics
is online at [Link] Abstract
[Link]
Collectively, various tandem and interspersed repetitive sequences make
014017
up approximately half the human genome, yet we have only begun to un-
Copyright © 2025 by the author(s). This work is
derstand the potential functions of “junk” DNA. Here, we provide a brief
licensed under a Creative Commons Attribution 4.0
International License, which permits unrestricted overview of various types of repeats, but a full treatment of the repeat
use, distribution, and reproduction in any medium, genome (repeatome) is beyond the scope of any review. Hence, we focus
provided the original author and source are credited.
primarily on less established functions of a few major repeat classes, includ-
See credit lines of images or other third-party
material in this article for license information. ing pericentromeric satellites and abundant degenerate interspersed repeats,
short interspersed nuclear elements (Alu), and long interspersed nuclear ele-
ments (L1). A theme developed throughout is how sequence organization in
the human karyotype provides insights into potential functions within nu-
clear structure. For example, millions of small tandem major satellite repeats
can form bodies that sequester nuclear factors, or the segmental organiza-
tion of interspersed repeats may underpin the nuclear compartmentalization
of heterochromatin and euchromatin. Decoding the vast repeatome is an
exciting frontier being enabled by recent technological advancements. How-
ever, identifying the extent of meaningful information in repeats will likely
require concepts that go well beyond impacts for individual genes, to new
ways to identify and interpret broad patterns of genome-wide organization
and nucleus-wide regulation.

45
1. INTRODUCTION
The Human Genome Project initially focused on sequencing the ∼20,000 protein-coding genes,
and there was debate as to whether the rest of the genome, riddled with repetitive sequences,
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

was worthy of the cost and time to interrogate it. However, we now know that polymorphisms
of interest are often in repeat-rich noncoding regions. Repetitive sequences have been notori-
ously difficult to investigate and map and thus have often been screened out of analysis using
RepeatMasker ([Link] or similar programs. Recent advances in long-
read sequencing technology make it possible to read through and precisely map sequences in
large repetitive regions, allowing comprehensive chromosome sequencing in the Telomere-to-
Telomere (T2T) project (83, 127). This has produced the first full, contiguous sequencing of the
repeatome, including large satellite regions, revealing greater structural complexity than previ-
ously anticipated. Multiple studies have begun to extend this technology to create pangenome
reference sequences that capture polymorphic differences in various repeats in populations (re-
viewed in 155). As we obtain a more complete description of the human genome’s repeat content
and its variations, a major challenge will be to assess the potential impacts of variation in different
types of repeats, and to understand the extent to which various types of repetitive “junk” play any
functional roles.
To undertake a review on the emerging functions of the various abundant repeats in the hu-
man genome is a timely but dauntingly large task. A full accounting of human repeats is beyond
the scope of any review, and we apologize that we cannot fully represent the huge literature of
related work. Here, we present a conceptual overview of emerging functions, focusing largely on
a few of the most abundant repeats with less established functions. Gene regulation has been most
studied in terms of local sequence effects on individual genes, and repeat elements often func-
tion at that level. However, we will convey our perspective that highly abundant repeats may also
function more collectively, to influence regional genome regulation in nuclei. This may best be
understood through the lens of human genome organization on chromosomes and how it relates
to compartmentalized genome regulation within complex nuclear structure.

1.1. Half the Human Genome Is Composed of Different


Types of Repeat Sequences
More than 50 years ago, DNA reannealing experiments discovered that large portions of the
genomes of higher organisms are composed of highly or moderately repeated sequences (20).
The classical C0 t curves (Figure 1a) show that abundant repeats reanneal at the lowest concen-
tration and time (C0 t-1), while unique or low-copy sequences reanneal more slowly. Subsequent
studies indicated that much of this repetitive DNA was interspersed throughout the genomes of
many organisms, although the proportion of the genome occupied by repeats can vary widely (41,
69). More than half a century ago, Britten & Davidson (19) theorized that these sequences may
function to regulate the genome. A few years earlier, cesium chloride density gradient centrifu-
gation studies of DNA from several species had resolved a primary band of DNA and a smaller
satellite band (97) (Figure 1b). This DNA satellite was subsequently shown to comprise large
tandem arrays of short repeated sequences (165) that are AT rich (37) and localized mostly at or
near chromosome centromeres (89).
Notably, half the human genome is composed of various repetitive sequences (Figure 1c),
which can be categorized as one of two major types: tandem repeats, which localize to specific
chromosomal sites, or interspersed repeats, which distribute widely through chromosomes, mostly
as single repeat units (Figure 1d). As discussed in Section 3, most interspersed repeats were derived
from mobile transposable elements (TEs) that invaded the human genome, but why degenerate
forms of TEs remain so abundant is an unsolved question in genome biology.
46 Hall • Smith • Lawrence
a c Other unique
sequences L1 (16.77%)
(18.55%) Other LINEs
0 + (3.90%)
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

+ Protein-coding
10 + sequences
(1.50%) SVAs
+
Fraction adsorbed

20 (0.15%)
Low-copy/unique
30 sequences Alu
40 (10.09%)
50
Repetitive Other
DNA Introns SINEs
60 + (26.00%) (2.68%)
70
+ LTRs (8.84%)
80
90 + Other (0.36%)
DNA transposons
100 Housekeeping RNA genes
10–3 10–2 10–1 100 101 102 103 104 (rRNA, tRNA, etc.) (0.11%) Satellites (3.58%)
Simple (4.92%)
C0t (mol × s/L) repeats
(2.54%)

b d LINEs
Non-LTRs
Class I SINEs
retrotransposons
Transposable
elements LTRs HERVs
Interspersed Class II
Density transposons
gradient Satellite band Microsatellites DNA
transposons
Main DNA band Repeats Minisatellites

Macrosatellites
Tandem
Centromeric
satellites
Pericentromeric
satellites

Telomeres

e f
Euchromatin Heterochromatin
Specific gene/
mRNA
Centromeres
1 2 3 4 5 6

Cajal bodies
Nucleolus
7 8 9 10 11 12

rDNA U2
genes
Specific gene/
13 14 15 16 17 18 mRNA
Telomeres Nuclear
Centromeres Nuclear lamina
rDNA speckles
HSat2
HSat3
19 20 21 22 X
(Caption appears on following page)

[Link] • Repeat Genome Functions in Nuclear Structure 47


Figure 1 (Figure appears on preceding page)
(a) Britten & Kohne’s (20) graph of the renaturation kinetics of calf thymus DNA (blue circles and triangles), showing ∼40% rapidly
reannealing repetitive sequences and ∼50–60% more unique sequences. Escherichia coli DNA (green plus signs) lacks the highly repetitive
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

fraction. Panel adapted with permission from Reference 20. (b) Density gradient centrifugation separating a smaller satellite band from
the main band of genomic DNA. (c) Pie chart showing the relative abundances of repeat types in the human genome. Data are from
Reference 83. (d) Categories of repeat types. Note that not all repeat types in each category are shown here. (e) Karyotype showing
alternating light and dark bands on a G-banded mitotic chromosome spread, with the locations of several repeat sequence types
indicated. Panel adapted with permission from Reference 145. ( f ) Illustration depicting the organization of the interphase nucleus. The
genome has a compartmentalized architecture within the nucleus and is organized further relative to non-membrane-bound
substructures rich in RNA metabolic factors, as indicated. Panel adapted with permission from Reference 145. Abbreviations: HERV,
human endogenous retrovirus; HSat, human satellite; L1, long interspersed nuclear element 1; LINE, long interspersed nuclear
element; LTR, long terminal repeat; SINE, short interspersed nuclear element; SVA, SINE-VNTR-Alu; VNTR, variable number
tandem repeat.

While the potential biological significance for the bulk of repetitive sequences is not known,
numerous studies have shown that a specific repeat sequence, typically near or in a protein-coding
gene, can impact the function of that gene through numerous different mechanisms, which can
be mediated by DNA or RNA. Changes in the location or copy number of a repeat can have
deleterious effects and contribute to disease or can be co-opted during evolution to contribute to
normal gene function. While there are now many examples of a repeat sequence being co-opted to
impact local gene function, they do not necessarily indicate whether the bulk of highly abundant
degenerate repeats contribute to genome function more broadly or if they are just an evolutionary
vestige. We discuss here less established concepts for how certain repeat types, present in enor-
mous numbers (hundreds of thousands to a million), may contribute to the broader regulation of
the genome within nuclear structure.
Identifying and understanding novel mechanisms may require different conceptual approaches
that go beyond the better-known molecular mechanisms that regulate individual genes. Adding
to the challenge, certain repeat types may function only transiently at particular stages of early
development, in response to stress, or in specific disease states, examples of which are mentioned
throughout this review. We consider these functions from the perspective of repeat genome orga-
nization in the human karyotype, which we suggest can provide insight into potentially broader
collective roles of abundant repeats in genome regulation.

1.2. Location, Location, Location: Connecting Chromosomal Organization


and Nuclear Genome Function
As the title suggests, this review discusses repeat sequences with “a view from the human kary-
otype,” with the underlying premise that the higher-order chromosomal distribution of repetitive
sequences (and genes) has been shaped through evolution to facilitate function. While the lo-
cations of telomeric or centromeric repeats clearly reflect their specific roles in chromosome
structure (discussed below), the linear sequence organization on chromosomes can also relate to
their function within the nucleus. The most straightforward demonstration of this is that tan-
dem copies of rDNA genes (and associated satellite repeats) occupy the small short arms of all
five acrocentric human chromosomes (Figure 1e). In our view, this singular organization evolved
to facilitate the formation of the nucleolus, a factory that promotes highly efficient rRNA tran-
scription, processing, and assembly of the numerous components required to produce abundant
ribosomes. Several other examples illustrate this principle on a smaller scale; for example, the
clustering of tandemly repeated U2 small nuclear RNA genes facilitates their association with the
nuclear Cajal body involved in the assembly of small nuclear ribonucleoproteins. Hence, the or-
ganization of DNA sequences, and often the RNAs they produce, can nucleate efficient nuclear
hubs for complex functions, which we previously referred to as the karyotype-to-hub hypothesis

48 Hall • Smith • Lawrence


(145). This principle is also relevant here as we consider the complexity and puzzling abundance
of nongenic repetitive sequences throughout the genome.
The human genome is full of larger cytogenetic patterns, revealing sequence organization so
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

pronounced that it is evident from simple staining and light microscopy. In addition to numerous
and often huge pericentric satellites, staining shows a pattern of 400–600 alternating light and
dark Giemsa-stained bands (Figure 1e), which correspond largely to regions with differences in
gene density, short interspersed nuclear elements (SINEs) versus long interspersed nuclear ele-
ments (LINEs), and GC versus AT content. What might be the functional significance of these
cytological-scale differences in linear genome organization? We suggest that a full understanding
will require a perspective on the complex substructure of the interphase nucleus, including the
compartmentalization of euchromatin and heterochromatin into large distinct nuclear regions,
and in cell type–specific patterns (Figure 1f ). The large euchromatin compartment, which is typ-
ically more internal in the nucleus, is punctuated by ∼10–20 discrete nuclear speckles (also known
as SC35 domains) that are concentrated with a host of RNA metabolic factors. This will become
important in Section 4 when we consider the large segmental organization of the genome, as re-
flected in chromosome bands, with differences in the density and types of genes as well as repetitive
sequences.
As alluded to above, the organization of telomere repeats (Figures 1e and 2a) clearly reflects
their function, to cap each chromosome end and protect it from fusing with other chromosomes
(see 28). However, telomere biology also illustrates that a given repeat sequence can have more
than one function and that the study of repeats can reveal unanticipated and fundamentally impor-
tant biology. The discovery that attrition of the telomere array is essentially a cellular aging clock
fueled numerous important discoveries in developmental biology and disease, particularly cancer
(reviewed in 6). Arrays of the telomere repeat TTAGGG are several kilobases in newborns and
are protected by the shelterin complex (reviewed in 43), but telomeres shorten progressively with
each somatic cell division; when they reach a critical length, a DNA damage response then trig-
gers cell senescence. In pluripotent cells, telomere length is maintained by telomerase, an enzyme
largely absent in differentiated cells, leading to telomere shortening and cell senescence (reviewed
in 28, 55). This example affirms the compelling prospects to uncover important new biology by
mining for meaningful information in the complex dark matter of human repetitive sequences.

2. EMERGING ROLES OF PERICENTRIC SATELLITES


IN NUCLEAR STRUCTURE
Tandem repeats vary greatly in terms of size of the repeat unit and length of the array. Categorized
largely by array size, satellites are the largest, followed by macro-, mini-, and microsatellites, with
some overlap between category definitions (see the sidebar titled Categories of Human Tandem
Repeats). We briefly discuss these smaller satellite types before focusing on the very large major
satellite arrays, especially those without a known or established function.

2.1. Many Mini- and Microsatellite Repeats Can Impact the Functions
of Specific Disease-Associated Genes in Cis
Diverse short tandem repeats (STRs) are present at many loci across our genomes, and changes
in individual tandem repeats can cause dysfunction in specific disease-associated genes. Approxi-
mately 50 monogenic diseases have been linked to the expansion or contraction of STRs (mostly
triplet repeats) in or near disease-causing genes, which can produce gain- or loss-of-function
effects on normal genes. These are primarily neurological diseases such as Huntington disease
(CAG), fragile X syndrome (CGG), myotonic dystrophy (CTG), Friedreich ataxia (GAA),
spinocerebellar ataxia (CAG), and amyotrophic lateral sclerosis (reviewed in 47).
[Link] • Repeat Genome Functions in Nuclear Structure 49
αSat, monomeric/divergent Other
a Telomere b Segmental HSat1 αSat HOR, inactive HSat2 satellite
duplications βSat αSat HOR, active HSat3
Centromere
DNA p arm q arm
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

Repeat
structure

HSat3B5
D9Z4
Chromosome 9

D7Z2

D7Z1
Large pericentromere
Chromosome 7 Moderate pericentromere

DXZ1
X chromosome Little or no pericentromere

Control Heat stress 3h 6h 24 h


c

DNA
YTHDC1 HSF1
Normal condition Heat stress Recovery from heat stress

Transcription-related nSB proteins


factors

HSat3

Heterochromatinized Euchromatinization and nSB assembly nSB remodeling


HSat3 regions HSat3 induction

Breast tumor
d e
Normal cell Inactive genes Cancer cell

HSat2 at 1q12 Other HSat2 sites

Demethylation

Demethylated
HSat2 at 1q12
sequesters PRC1
in CAP body Loss of ubH2A
Stochastic
expression
of activated HSat2 RNA
1q12 remains genes DNA
transcriptionally
PRC1 and repressed HSat2 RNA PRC1 and
MeCP2 evenly sequesters MeCP2 MeCP2 in
distributed in CAST bodies nuclear bodies
Demethylated
DNA Methylated DNA PRC1
Loss of UbH2A HSat2 RNA MeCP2
(Caption appears on following page)

50 Hall • Smith • Lawrence


Figure 2 (Figure appears on preceding page)
(a) Digital image of human telomeres and centromeres on metaphase chromosomes detected by FISH using labeled oligonucleotide
probes for telomeric (green) and centromeric (red) sequences and DAPI DNA dye (blue). Image provided by the laboratory of Dr. Jerry
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

W. Shay. (b) Schematic of a generalized human peri/centromeric region. The amount of pericentromeric sequence varies greatly among
different chromosomes. Panel adapted with permission from Reference 3. (c) nSB formation. (Top) YTHDC1 (green) and HSF1 (red)
fluorescent immunostaining and DAPI DNA staining (blue) in control cells, recruitment of HSF1 upon heat stress, sequestration of
YTHDC1 at 3 and 6 h during recovery, and return to normal after 24 h. (Bottom) Diagram of HSat3 expression and nSB formation.
HSat3 repeats, which are normally heterochromatin, become expressed upon heat stress. HSat3 RNA recruits specific proteins,
assembling nSBs, which are subsequently remodeled through recruitment of other factors. Top subpanel adapted from Reference 156
(CC BY 4.0); bottom subpanel adapted from Reference 124 (CC BY 4.0). (d) Schematic of CAP and CAST body formation by HSat2.
In many tumors, DNA demethylation triggers HSat2 DNA and RNA molecular sponges, causing further epigenetic dysfunction.
Sequestration of PRC1 at the demethylated 1q12 megasatellite forms CAP bodies and reduces the repressive ubH2A modification at
other HSat2 loci, which express RNA that sequesters MeCP2 in CAST bodies. Panel adapted with permission from Reference 72
(CC BY-NC-ND 4.0). (e) Image showing HSat2 RNA (green) forming CAST bodies in breast tumor cells (with DAPI-stained nuclear
DNA shown in blue). Panel adapted with permission from Reference 72 (CC BY-NC-ND 4.0). Abbreviations: αSat, alpha satellite;
βSat, beta satellite; CAP, cancer-associated Polycomb; CAST, cancer-associated satellite transcript; DAPI, 4′ ,6-diamidino-2-
phenylindole; FISH, fluorescence in situ hybridization; HOR, higher-order repeat; HSat, human satellite; nSB, nuclear stress body;
ubH2A, ubiquitinated histone H2A.

These repeats are also thought to play a variety of roles in normal gene function (reviewed
in 7), including operating as part of gene products [e.g., in coding exons (151) or noncoding
RNA (ncRNA) functional domains (21)], by influencing local chromatin structure and tran-
scription [e.g., nucleosome spacing, CpG methylation, transcription factor (TF) binding sites,
transcription start sites, and enhancers] and acting within untranslated regions or introns to
modulate transcription, translation, and alternative splicing. These highly variable arrays provide a

CATEGORIES OF HUMAN TANDEM REPEATS


■ Satellites: Satellites are generally very large arrays (megabases) of repeated units that can range from 5 bp
to several hundred base pairs and collectively comprise 6.2% of the human genome. There are two main
types: centromeric satellites, which make up the main functional centromere of each chromosome, and
pericentromeric satellites, which are adjacent to some but not all human centromeres (reviewed in 2, 121).
■ Macrosatellites: Microsatellites have a repeat unit size of >100 bp (typically ∼1–6 kb), in arrays ranging from
a few kilobases to hundreds of kilobases in length; each is usually found in only one or two loci in the genome.
Although their sequences are unrelated, they tend to be CpG rich, often produce noncoding and/or coding
RNAs, and tend to be regulated by DNA methylation (reviewed in 52).
■ Minisatellites: Also called variable number tandem repeats (VNTRs), minisatellites are typically defined as
having repeat units of 6–100 bp (158) and highly variable array lengths (typically from 0.5 kb to several
kilobases). There are approximately 1,000 minisatellites in the genome, located mostly near the ends of
chromosomes.
■ Microsatellites: Also called short tandem repeats (STRs), microsatellites have repeat units of 1–6 bp repeated
∼6–30 times; ∼500,000 of them are scattered across the genome, constituting ∼2.5% of the human genome
(66, 83). Microsatellite arrays are highly polymorphic in length, primarily due to a replication error called
strand slippage, as well as misalignment at meiosis and unequal crossover. Changes in the copy number of
STRs at specific gene loci have been linked to many diseases, as discussed in Section 2.1. Their abundance
and variability have also made them important tools in molecular diagnostics, population studies, and foren-
sics (reviewed in 66). In Section 4.5 we make a distinction between localized STRs and interspersed simple
sequence repeats (i-SSRs), which are extremely high-copy-number repeat “words” dispersed throughout the
whole genome.

[Link] • Repeat Genome Functions in Nuclear Structure 51


larger polymorphic range than the more binary single-nucleotide polymorphisms, and numer-
ous studies implicate these repeats in many common human disorders, including neurological
and neurodegenerative disorders, cardiovascular disease, diabetes, and cancer (reviewed in 47, 77,
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

117). Therefore, STRs could contribute to the missing heritability in multifactorial conditions
and in normal phenotypic variation. And with recent advances in sequencing technology, their
true variation is now being explored (e.g., 137).
We highlight just one of the first triplet repeat disorders discovered, which illustrates a theme
developed further below for major satellites: the capacity of very abundant small repeats to bind
and sequester regulatory factors. Myotonic dystrophy type 1 results from a large expansion of
CTG triplet repeats in the 3′ untranslated region of the DMPK gene. This causes the DMPK
mRNA containing the repeats to accumulate to high levels in the nucleus, forming ribonucleopro-
tein aggregates that sequester an important splicing regulator, MBNL (muscleblind-like). MBNL
levels throughout the nucleoplasm drop sharply as a consequence (144, 167), which impairs alter-
native splicing of pre-mRNAs for many other genes (reviewed in 111). The concept that highly
abundant repeats can bind and impact the distribution of specific nuclear factors will be important
as we consider the function of the much larger tandem repeats, satellites.

2.2. Large Major Satellite Arrays of Small Repeats at Specific Loci:


Not-So-Constitutive Heterochromatin
The much larger major satellites are very different in form and function than the smaller tan-
dem repeats discussed above, as these huge arrays are located predominantly at one location: at or
near the centromeres of chromosomes. The three major satellites we discuss here—alpha satel-
lite (αSat), human satellite 2 (HSat2), and HSat3—are multi-megabase arrays comprising tens
or hundreds of thousands of small repeats at a single locus. Historically thought to be constitu-
tively silent heterochromatin, large satellites are now recognized to be transiently expressed in
different contexts during early embryonic development, cell cycle stages, specific diseases (such
as cancer), or changes in cell state such as stress (discussed in Section 2.4.1). The transient and
complex nature of satellite expression patterns will make it more challenging to fully investigate
their biological functions, but results to date make clear that they can serve important functions
at the DNA and/or RNA level.
Why did the genome evolve to accumulate many thousands or millions of tandem copies
of a small sequence in singular locations? All of these large satellites are in or adjacent to the
centromere, and we begin by discussing the more well-established role for αSat in centromere
function. The functions of pericentric satellites (HSat2 and HSat3) are less clear: While they may
play a structural role related to centromere function, we highlight emerging evidence that they
can also play a role in global genome regulation in nuclei.

2.3. Centromeric αSat Repeats Localize Proteins to Form the Kinetochore


of Segregating Chromosomes
Centromeres (Figure 2a,b) are composed of a 171-bp αSat repeat unit in a single 2–5-Mb array
on each chromosome, producing a structure essential for kinetochore assembly and function dur-
ing mitosis and meiosis. Centromeric satellite DNA has both active and inactive histone marks
and is transcribed at low levels in human cells. The transcripts form DNA:RNA hybrids that sta-
bilize an RNA:protein structure at the locus that facilitates recruitment and stabilization of most
centromeric proteins [centromere proteins (CENPs), passenger complex, etc.] required for cen-
tromere function (reviewed in 172). Preventing centromere transcription leads to gradual loss of
these centromeric proteins and genomic instability (reviewed in 36).

52 Hall • Smith • Lawrence


The accumulation of CENP-A (a histone H3 variant) in nucleosomes uniquely distinguishes
the centromere proper from the rest of the genome, and this chromatin serves as the platform for
kinetochore assembly, an enormous complex that binds spindle microtubules during cell division
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

(reviewed in 91). The accumulation of ∼100 different proteins on αSat centromeric DNA illus-
trates how high-copy repeats organized into arrays serve to concentrate protein components to
build a structure that functions at that site.
αSat array size and sequence polymorphisms have been associated with defective centromere
architecture and aneuploidies, and polymorphic satellite array size can vary between homologs (re-
viewed in 121, 152). This suggests that these polymorphic differences may play important roles
in human health, but until recently, human centromeres were almost entirely absent from the
genome build, hindering their study. The T2T gapless assembly published in 2022 for the first
time includes all human centromeric sequences, and the inclusion of diverse populations has re-
vealed more variability than expected in αSat sequences, especially among people of recent African
origin (3).

2.4. Diverse Functions for the Huge Pericentric Satellites: HSat2 and HSat3
Pericentric satellites adjacent to the centromere (Figure 2b) were nearly absent from the refer-
ence human genome until recently (2), making them significantly understudied. The two most
abundant are HSat2 and HSat3, which total 28.7 and 47.6 Mb, respectively, and are found on nu-
merous, but not all, human chromosomes. HSat3 is derived from a pentameric repeat, (CATTC)n ,
and the HSat2 repeat is an ∼26-bp degenerate sequence derived from the HSat3 pentamer. To-
gether, they constitute the largest contiguous satellite arrays in the human genome, including an
∼28-Mb HSat3 array on chromosome 9 and the two largest HSat2 arrays, on chromosomes 1 and
16, which are approximately half that size (∼14 Mb) (2) (Figure 1e).
Pericentric satellites are generally silent in most normal cells, with the exception of testis and
brain, and their heterochromatic nature may help stabilize the centromere (reviewed in 60). How-
ever, not all human chromosomes have pericentric satellites (Figure 2b), which indicates that they
are not necessarily required for normal centromere function. Nevertheless, aberrations in their
heterochromatic state or expression have been associated with mitotic defects in spindle attach-
ments and sister chromatid cohesion, as well as increased DNA damage in S-phase due to blocked
replication over pericentric DNA:RNA hybrids (reviewed in 146).
Interestingly, however, pericentric satellites harbor promoter elements that can regulate tran-
scription by RNA polymerase II (RNAPII) or RNAPIII, and pericentric satellite expression is
common during embryogenesis. In fact, many different human satellite families are expressed in
complex patterns during early embryogenesis that appear to be highly regulated (reviewed in 120,
146). This implies directed regulation of individual satellite arrays during specific windows of em-
bryonic development, for currently unknown reasons. One possibility is that pericentric satellite
expression early in embryogenesis plays a role in nucleating the formation of initial heterochro-
matic compartments with this unique chromatin (reviewed in 133). However, this has yet to be
fully explored, as these regions are not currently included in the genome maps for most species
and have only recently been added to the human genome build.
Another possible function is that both HSat2 and HSat3 can impart global gene regulation
through their capacity to act as a cytological-scale molecular sponge. The collective evidence
summarized below suggests that these exceptionally high-copy pericentric arrays, containing
repeat units with protein-binding potential, have extraordinary capacity to amass and cytolog-
ically sequester regulatory factors at both the DNA and RNA levels, thereby modifying their
accessibility on a genome-wide scale. For example, a 14-Mb array of a 26-nt repeat unit will
contain ∼500,000 copies, while a 5-nt sequence could be repeated millions of times in a single

[Link] • Repeat Genome Functions in Nuclear Structure 53


multi-megabase location. Hence, there is a remarkable potential to congregate or sequester
factors at one location within nuclear structure. Interestingly, centric satellite repeats function
similarly in various species, yet the satellite sequence itself is not conserved (e.g., between human
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

and mouse). This illustrates that the function of repeats is often less stringently tied to primary
sequence than it is for protein-coding genes.

2.4.1. HSat3 DNA and RNA: the nuclear stress sponge. The earliest and most developed
evidence of a human satellite functioning as a sponge is for the very large HSat3 array on chro-
mosome 9 (9q12) and more recently recognized on the Y chromosome. These arrays act at both
the DNA and RNA levels to regulate cell homoeostasis during numerous types of cell stress (e.g.,
heat shock, osmotic or oxidative stress, and UV radiation). Cell stress triggers a series of steps in
which different factors are sequentially bound and released from HSat3 DNA or RNA during the
stress response, including a prolonged recovery process (reviewed in 61, 124) (Figure 2c). At stress
onset, HSF1 (heat shock transcription factor 1) is expressed, and the HSF1 protein localizes to the
normally silent HSat3 loci on chromosome 9 and the Y chromosome, along with several other
TFs and chromatin-remodeling factors. HSF1 activates HSat3 transcription via RNAPII, and the
HSat3 ncRNA transcripts accumulate at the locus, forming ribonucleoprotein bodies called nu-
clear stress bodies. Nuclear stress bodies sequester many different RNA metabolic factors, leading
to global suppression of transcription and translation, until the stress is resolved. During stress re-
covery, these components are released to the nucleoplasm to reactivate the genome in a highly
regulated manner. Nuclear stress bodies remain during the prolonged stress recovery period and
sequester other factors to help reverse the process, suggesting that the same bodies can dynam-
ically change their properties and function. In the final stage of stress recovery, the remaining
HSat3 transcripts recruit repressive factors to re-silence the HSat3 loci, making the sponge inert
once again.
The ability of repetitive RNA to nucleate phased domains confers a second means to affect
genome regulation broadly: by concentrating specific factors together in a reaction crucible that
accelerates biochemical reactions. HSat3 nuclear stress bodies also behave this way during cell
stress (reviewed in 124), illustrating the versatility of satellite RNA bodies in broadly regulating
the genome.

2.4.2. HSat2 DNA and RNA: the nuclear disease sponge. The large HSat2 arrays on chro-
mosomes 1 and 16 are two of the most prominent but poorly studied features of the human
genome. Recent studies have uncovered unanticipated biology of HSat2 satellites that points to
effects mediated by DNA demethylation and RNA expression, but with complex differences be-
tween HSat2 loci on different chromosomes. Translocations and duplications of the large HSat2
array at 1q12 are among the most frequent aberrations in cancers (reviewed in 70), and global DNA
demethylation is also common in human cancers [and in ICF (immunodeficiency, centromeric
region instability, and facial anomalies) syndrome (160)], with HSat2 at 1q12 being especially
sensitive to demethylation (53).
This global demethylation appears to trigger HSat2 at 1q12 to act as a molecular sponge
(72), which further alters the epigenetic state of the cell. DNA demethylation causes PRC1
(Polycomb repressive complex 1), which normally maintains the repressive ubiquitinated his-
tone H2A (UbH2A) mark at target gene loci across the genome, to accumulate over the 1q12
locus, forming large cancer-associated Polycomb (CAP) bodies (72) (Figure 2d). Although sev-
eral human chromosomes have smaller pericentromeric HSat2 arrays, this is a locus-specific
(1q12) and protein-specific (PRC1 but not PRC2) response to DNA demethylation in human
cells, while in mouse, both PRC1 and PRC2 (which trimethylates H3K27) are sequestered to
all major satellites in pericentromeres upon demethylation (33). This difference likely reflects

54 Hall • Smith • Lawrence


the greater chromosome-specific sequence diversity in human satellites compared with mouse
pericentric satellites (84, 162) and suggests that humans evolved locus-specific satellite function.
The functional sequestration of epigenetic factors to large pericentromeric arrays can have
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

downstream consequences for global transcriptional regulation, which results in aberrant ex-
pression across the genome (33, 72). Importantly, this includes small HSat2 arrays on other
chromosomes, which become derepressed and aberrantly expressed (72) (Figure 2d). In fact, nu-
merous studies have shown that normally silent HSat2 repeats are commonly overexpressed in
cancers, more frequently than any other satellite (13, 72, 96), as well as during viral infection,
senescence, and DNA damage and in diseases like facioscapulohumeral muscular dystrophy (e.g.,
125, 141).
Aberrant expression from HSat2 loci then compounds the epigenetic dysregulation in these
cells (Figure 2d) by sequestering additional regulatory factors into large ribonucleoprotein bodies.
These HSat2 RNA bodies are prominent hallmarks of many tumors (Figure 2e), detected in ap-
proximately half of 34 diverse tumors examined, and sequester MeCP2 (methyl-CpG binding pro-
tein 2) (72, 102). These were initially termed cancer-associated satellite transcript (CAST) bodies;
however, HSat2 RNA bodies are also seen in other disease contexts, where we call them satellite
transcript (SATT) bodies, and sequester different regulatory factors, including CTCF (CCCTC-
binding factor) (122), EIF4A3 (eukaryotic translation initiation factor 4A3), and ADAR1
(adenosine deaminase RNA 1) (141), leading to further dysregulation of cell homeostasis and gene
expression. Thus, HSat2, like HSat3, can form both DNA and RNA molecular sponges that impact
genome-wide access to important regulatory factors and broadly affect gene expression.
2.4.3. HSat2 expression in development and disease and potential regulatory effects on
gene pathways. Aberrant expression of satellites in disease may not only be a consequence of
misregulation but also directly contribute to effects on specific downstream pathways. There is
some evidence to suggest that cancer cells or viruses co-opt HSat2 expression to impact pathways
that confer a growth advantage. For example, tumors appear to select for expression of specific
satellites (72, 148), and many human viruses use TFs to specifically activate HSat2 loci (126). The
presence of these satellite RNAs (particularly HSat2) is associated with gene expression changes
that confer reduced immune response, changes in cellular motility, or changes in protein stabil-
ity and localization (e.g., 126, 134, 148). And some work has also shown a direct link between
the presence of the HSat2 ncRNA and aberrant expression from specific pathways (loss of the
RNA prevented the effect) (126). This suggests that HSat2 RNA itself altered the regulation of
specific gene pathways via an unknown mechanism, which may well be through sequestration of
their regulatory factors. However, most studies do not look for expression from the repeatome or
sequestration to RNA bodies, so it is unclear whether SATT bodies are responsible for the gene
expression changes observed.
Most human satellite families are enriched in a wide variety of satellite-specific TF binding
sites, including those that regulate conserved signaling pathways (62, 162). Since expression of
different satellite families appears to be highly regulated throughout embryogenesis, and global
hypomethylation of satellites is also a normal hallmark of gametes, preimplantation embryos, and
extraembryonic tissues (reviewed in 163), it is tempting to speculate that satellites may act as DNA
or RNA molecular sponges at important transition points during normal development. This may
dynamically regulate global genomic access to specific regulatory factors during developmental
transitions.
The likelihood that satellites can act as DNA or RNA sponges during embryogenesis is
supported by findings that mouse pericentromeres (which form chromocenters) can functionally
sequester specific TFs during mouse embryogenesis (110). Additionally, the DUX4 (double ho-
meobox 4) TF is expressed during specific windows of human embryogenesis [and by many human

[Link] • Repeat Genome Functions in Nuclear Structure 55


cancers and viruses (reviewed in 123)] and is required for proper maturation of pre- and postim-
plantation embryos (141, 161). DUX4 induces expression of HSat2 (and other repeats) during
these same developmental windows, with HSat2 SATT body formation and sequestration of key
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

regulators.
The fascinating story of DUX4 began with the discovery of its abnormal activation in fa-
cioscapulohumeral muscular dystrophy (50) and illustrates the importance of including the
repeatome in both transcriptomic and molecular cytology studies (reviewed in 123). The
macrosatellite D4Z4, on the subtelomere of chromosome 4q, is heterochromatic and silent in
most adult cell types. However, reducing the copy number of the 3.3-kb repeat in the D4Z4 array
to <10 triggers loss of heterochromatic repression and aberrant DUX4 expression, which is con-
sidered the primary cause of muscle degeneration (71). DUX4, which normally regulates HSat2
expression during development, induces aberrant HSat2 RNA in facioscapulohumeral muscular
dystrophy muscle. HSat2 SATT bodies in cell nuclei sequester important regulatory factors that
affect RNA stability, splicing, and translation. Hence, future studies will need to consider what
downstream consequences are directly due to the DUX4 TF or, alternatively, might be due to
effects of sequestration of regulatory factors by HSat2.
This example illustrates how a macrosatellite can act locally to regulate the epigenetic state of a
locus (4q35) that encodes an important regulatory gene (DUX4), which in turn normally regulates
embryonic expression of a major satellite (HSat2), potentially affecting nuclear regulatory factors
more globally.

3. ABUNDANT INTERSPERSED REPEATS DERIVED FROM MOBILE


TRANSPOSABLE ELEMENTS
The largest portion of the repeatome consists of interspersed repeats derived from TEs that long
ago invaded the human genome and influenced its evolution. TEs are mobile DNA sequences
capable of transposition and integration into the genome and can be divided into two broad
categories based on their mechanism of transposition. Class I TEs are retrotransposons, which
mobilize by transcription into an RNA intermediate and then reverse transcription into a new
genomic location, while class II TEs are DNA transposons that move as DNA via a variety of
cut-and-paste mechanisms (reviewed in 166).
We are relatively early in uncovering the full impact that TEs have had on human disease and
genome evolution. Intact mobile elements spread broadly and gave rise to copious degenerate re-
peats that can be co-opted for the regulation and function of individual genes (17, 35). Only a very
small number of TEs remain intact and capable of transposition, but abundant degenerate TEs
that can no longer hop remain; for clarity, we refer to the latter as TE-derived sequences (TEDS).
TEDS derived from LINE1 (L1) and SINEs (Alu) are the largest class of these interspersed repeats
(Figure 1c) and are our main focus here. The challenge is to understand the potential functional
raison d’être for millions of these degenerate TEs and to consider how their higher-level genomic
organization may relate to their functions.
Below, we briefly review some of the more abundant and/or active TEs and introduce the
various known ways that individual TEDS have been co-opted to contribute to the functions of
specific protein-coding genes. In Section 4, we discuss emerging evidence for the less established
roles of TEDS in broad genomic regulation. For more in-depth information on TEs, we refer
readers to several excellent recent reviews (14, 80, 150, 166, 170).

3.1. A Brief Introduction to the Major Types of Transposable Elements


LINEs are the largest family of retrotransposons and the largest family of repeats in the
genome, and L1 is the predominant LINE in humans. The intact L1 is ∼6 kb in length and

56 Hall • Smith • Lawrence


has a bidirectional RNAPII promoter (Figure 3a). Sense transcription produces ORF1p (open
reading frame 1 protein), an RNA-binding protein that interacts with L1 RNA, and ORF2p,
a protein with endonuclease and reverse transcriptase activities; both proteins are required for
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

retrotransposition (reviewed in 11). Additionally, a second promoter drives transcription in the


antisense direction, transcribing the primate-specific ORF0, which produces a small peptide
whose function is unknown but which may enhance L1 mobility. ORF0 contains splice donor
sites that allow it to form fusion proteins with exons from neighboring genes (46). There are
∼560,000 L1 sequences (the bulk of which are TEDS), comprising ∼17% of the human genome
(83) (Figure 1c). However, only ∼7,000 of these L1s have an intact promoter (94), and only
∼80–100 evolutionarily recent L1s are intact and capable of retrotransposition (22).
SINEs (reviewed in 170) are the most abundant TEs (and TEDS) in the human genome, and
Alu is the largest subfamily of SINEs, accounting for ∼10% of the genome (Figure 1c). Unlike
L1, Alu is a much smaller TE (∼280 bp) and is composed of left and right arms (derived from
7SL RNAs) separated by a poly(A) stretch, with the left arm containing weak RNAPIII promoters
(45) (Figure 3a). Alu and other SINEs do not encode reverse transcriptase and thus rely on L1
for retrotransposition (49). There are ∼1.15 million Alu sequences in the human genome (83),
but as with L1, the vast majority are TEDS, and only a fraction are still capable of retrotransposi-
tion (45, 99). Mammalian-wide interspersed repeat (MIR) elements are ancient mammalian SINEs
that number ∼600,000 (∼3%) in the human genome (170). Full-length MIRs are 260 bp long and
contain a tRNA-derived left arm, which contains RNAPIII promoters, a central SINE sequence,
and a LINE-derived right arm (170). MIR density has been correlated with tissue-specific gene
expression (87), and MIRs have been found to function as gene regulatory elements such as insu-
lators (164) and enhancers (86). While our focus here is on Alu and L1, below we briefly introduce
a few other, less abundant TEs and provide relevant reviews.
The hominid-specific SINE–variable number tandem repeat (VNTR)–Alus (SVAs) are the
youngest family of retrotransposons, with a complex composite structure including a hexamer
repeat (CCCTCT), an Alu-like sequence, a GC-rich VNTR, a SINE, and a poly(A) tail (76).
While there are relatively few (∼6,600) SVA elements in the genome (83), they can be active, but,
like Alu, they are nonautonomous, relying on L1 machinery for mobilization (136). SVA insertions
have been associated with numerous diseases (132), ∼60% are within 10 kb upstream of genes,
and evidence suggests they may be involved in gene regulation (135). Interestingly, SVAs are often
concentrated in clusters of KRAB–zinc finger protein genes (especially on chromosome 19), which
are themselves thought to be involved in the control of TEs (67).
Long terminal repeat retrotransposons are defined by the presence of flanking long terminal
repeat sequences, typically flanking two ORFs, gag and pol. They comprise ∼8% of the genome
and are primarily human endogenous retroviruses (HERVs), remnants of ancient retroviral in-
fections (reviewed in 150). Mutations have rendered most HERVs replication incompetent and
largely silent (88); however, the youngest family of HERV-K elements retains the capacity to
produce viral proteins and virus particles (82). HERV-K appears to be expressed in most normal
tissues (23) but is highly expressed in a variety of cancers, where it may promote tumor growth and
metastasis (48). The HERV-H family is also noteworthy because it is highly expressed in human
embryonic stem cells, and evidence indicates that it is essential for pluripotency (reviewed in 140).
Numerous families of DNA transposons collectively comprise ∼3.6% of the genome (83).
They typically have a transposase ORF flanked by terminal inverted repeats and can vary greatly
in size (166). Nonautonomous miniature inverted-repeat TEs do not encode a transposase and
rely on the machinery of other DNA transposons to mobilize (57). While no longer active (128),
DNA transposons have played a significant role in genome evolution in many organisms, including

[Link] • Repeat Genome Functions in Nuclear Structure 57


a ~6 kb c Transposition errors
and mutations
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

Poly(A)
L1 5’ UTR ORF1 EN
ORF2
RT C 3’ UTR tail
ORF0
Active TEs TEDS
~300 nt
Left arm Right arm New regulatory
Deletions elements
Alu A B AAA AAAAAA
Co-opted
Duplications New proteins for gene
function
b Active TEs
(0.02%)
Rearrangements New lncRNAs Genome
Disease evolution

TEDS Other sequences Aberrant gene


(45.98%) (54.00%) expression

Genome Shape Broader


instability chromatin genomic
architecture regulation

d Plasma cell Monocyte e DNA XIST RNA C0t-1 RNA

Untreated
nuclei

Extracted
nuclei

f g
Percentage of the genome

Genome DNA
15 Scaffold RNA
XIST RNA hCOT1 RNA
Soluble RNA
DAPI DNA DAPI DNA
10

0
HT1080 G3 GM11687 hybrid
L1

L2

IR
V

SR

ou r/
hA
Al

ER

gu e
M

s
/

bi th
LC

O
am

h
L1 DNA Alu DNA CpG DNA Late-replicating DNA

(Caption appears on following page)

58 Hall • Smith • Lawrence


Figure 3 (Figure appears on preceding page)
(a) Structure of L1 and Alu TEs. (b) Pie chart showing the proportions of TEDS, active TEs, and other sequences in the genome.
Degenerate TEDS are no longer mobile but vastly outnumber active TEs, which are a tiny fraction of the genome. (c) Diagram
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

showing the typically negative effects of TE transposition compared with the positive contributions of TEDS in the function of
individual genes, as well as their emerging broader role in nuclear genome architecture. (d) Line drawings of plasma cell and monocyte
nuclei illustrating how the organization of the condensed heterochromatic compartment and the more open euchromatin differs
between cell types. Panel adapted with permission from Reference 26. (e) Images of untreated and extracted nuclei showing that both
XIST RNA (red) and C0 t-1 RNA (green) remain localized and bound with the nuclear scaffold after nuclear extraction and removal of
histones and DNA. Panel adapted with permission from Reference 39. (f ) XIST RNA in HT1080 G3 cells (left) and C0 t-1 RNA in
GM11687 hybrid cells (right). Similar to how XIST RNA localizes to the inactive X chromosome territory (left), human C0 t-1 RNA
strictly localizes on the active human chromosome territory in hybrid cells (right). Panel adapted with permission from Reference 73.
(g) Graph of soluble or scaffold-associated repeat RNAs relative to their abundance in the genome. Repeats sequenced in nuclear RNA
are overwhelmingly associated with the insoluble nuclear scaffold. Panel adapted with permission from Reference 39. (h, left) Examples
of L1 and Alu distribution on several human mitotic chromosomes as detected by DNA FISH. (Right) The same chromosomes labeled
for CpG density and late replication. Left subpanel adapted with permission from Reference 74; right subpanel adapted with
permission from Reference 15. Abbreviations: C, cysteine-rich domain; DAPI, 4′ ,6-diamidino-2-phenylindole; EN, endonuclease; ERV,
endogenous retrovirus; FISH, fluorescence in situ hybridization; L1/2, long interspersed nuclear element 1/2; LC, low complexity;
lncRNA, long noncoding RNA; MIR, mammalian-wide interspersed repeat; ORF, open reading frame; RT, reverse transcriptase; SR,
simple repeats; TE, transposable element; TEDS, TE-derived sequences; UTR, untranslated region.

humans, by altering the structure or regulation of specific genes and triggering chromosomal
rearrangements (58).

3.2. Transposition Activity of Intact LINEs and SINEs and Its Consequences
Less than 0.05% of the millions of TE sequences remain intact and capable of transposition
(Figure 3b), and all of them are retrotransposons, including evolutionarily young families of intact
LINEs and SINEs. The rate of transposition in humans is low: Approximately 1 in every 17 births
carries a new TE integration (59). Since TE activation can have harmful effects, mobile TEs are
largely silenced via epigenetic mechanisms, including DNA methylation, histone modifications,
and silencing mediated by Piwi-interacting RNA and small interfering RNA (1, 34). While tightly
regulated in most somatic tissues, L1 expression and retrotransposition do occur at specific times
in development (154). Transposition of L1 is more frequent in early embryogenesis and occurs in
specialized cells such as spermatozoa and oocytes (65, 105). Low-level L1 and/or Alu activity may
also contribute to somatic mosaicism (92), particularly in the brain (reviewed in 16).

3.3. The Downside of Retrotransposition


Integration of a mobile TE in or near a gene will often disrupt normal gene structure or regula-
tion and can also lead to ectopic recombination, chromosomal rearrangements, duplications, or
deletions (35, 139) (Figure 3c). TE mobilization has been implicated in a number of neurodegen-
erative and neuropsychiatric disorders (34, 130), and TEs are frequently dysregulated in cancer
(24). L1 expression and retrotransposition can profoundly affect genome stability via DNA dam-
age and replication stress (5, 118) and have been reported to increase in aging and cell senescence,
driving interferon expression and inflammation (42).
The bidirectional L1 promoter (Figure 3a) can produce transcripts containing 5′ L1 antisense
sequences along with sequences of nearby genes. These chimeric transcripts are produced in sev-
eral cell types, potentially affecting up to 4% of human genes (40), and some are associated with
cancer (27, 95). L1 sequences also contain multiple putative splice sites that can cause aberrant
splicing, resulting in disease (e.g., 12, 168).
Like L1, Alu transposition can cause insertion mutations and recombination as well as affect
local gene expression and function through a variety of similar mechanisms (10). Alu activity has
also been implicated in a number of neurological and other diseases (103, 131).

[Link] • Repeat Genome Functions in Nuclear Structure 59


3.4. Some Transposable Element–Derived Sequences Have Been Co-opted
for a Role in Local Gene Function
While TE hopping can be deleterious to the host organism, over evolutionary time, mobile TEs
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

gave rise to millions of TEDS, which for poorly understood reasons have persisted to vastly
outnumber the active TEs (Figure 3b). Some TEDS have been domesticated for normal gene
functions, such as generating new functional regulatory elements, ncRNAs, and proteins (reviewed
in 56, 63) (Figure 3c). For instance, more than 20% of regulatory elements in the human genome
are TE derived, and more than 85% of these are primate specific (4, 51). Approximately 75%
of human genes have at least one Alu sequence, and examples of Alu regulating the function of
nearby genes are especially numerous. These include acting as cis-acting DNA regulatory ele-
ments (e.g., promoters, enhancers, insulators, or TF binding sites) or within mRNAs (in introns
or untranslated regions) to influence splicing, nuclear retention, and mRNA stability (reviewed in
170). A recent study showed that some enhancers may use RNA pairing to interact with specific
promoters and that almost 40% of these RNA interaction sites overlap Alu sequences (108). In ad-
dition, transcription of L1 sequences has enhancer functions that are essential to zygotic genome
activation in mouse embryos (107).
TE sequences are also a source for the evolution of new genes (Figure 3c). More than
80% of human long ncRNAs (lncRNAs) contain at least one TEDS, with TEDS compris-
ing ∼40% of lncRNA sequences (93). For example, the structural RNAs NEAT1 and XIST,
which are responsible for the formation of paraspeckles and X inactivation/Barr body forma-
tion, respectively, contain numerous repetitive sequences, some of which may derive from TEs,
and which serve as binding sites for proteins essential to their function (54, 169). In addi-
tion to lncRNAs, microRNAs and Piwi-interacting RNAs can also be derived from TEs. TEs
have also been exapted to create more than 100 new proteins. For instance, CENP-B, which
is involved in centromere formation, was derived from a DNA transposon. A variety of pro-
teins important in lymphocyte, placenta, and brain development are TE derived (reviewed in
56).
The studies cited above and numerous others have demonstrated that an interspersed repeat
sequence can contribute to the regulation or function of a nearby gene, or as part of the gene
itself. However, this does not necessarily attribute functionality to the sea of innumerable re-
peats interlaced through the whole genome. A major challenge remains to understand whether the
abundance of interspersed repeats serves some general genomic function or is mostly evolutionary
detritus. This question is the focus of the next section.

4. HIGHER-LEVEL ORGANIZATION OF INTERSPERSED REPEATS


LINKED TO NUCLEAR COMPARTMENTALIZATION
The collective ∼1.6 million Alu and L1 sequences far exceed the protein-coding portion of the
genome (∼100-fold), so might such abundant repetitive “junk” play a broader role in genome reg-
ulation, beyond the level of individual genes? Interspersed repetitive sequences are particularly
well-suited to propagate a pattern of chromatin folding across a larger region, using mechanisms
such as phase separation of repeat-binding proteins or a unique capacity to interact and form
unusual structures (G-quadruplexes, triplex DNA/RNA, etc.). Here, we consider emerging evi-
dence implicating repeats, including repeat-rich RNAs, in higher-level regulation of functional
nuclear architecture. In particular, we relate how differences in gene and repeat family distribu-
tions are organized across the human karyotype and how this relates to genome regulation within
compartmentalized nuclear structure.

60 Hall • Smith • Lawrence


4.1. The Nuclear Genome Segregates into Large Heterochromatin
and Euchromatin Compartments
The nuclear genome is packaged into two large, cytologically distinct compartments (Figures 1f
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

and 3d): condensed, inactive heterochromatin, mostly near the nuclear or nucleolar peripheries,
and open euchromatin, which occupies much of the interior nuclear regions in most cell types.
It may often be thought that the activity of individual genes explains the visible decondensation
evident throughout euchromatin, but it does not. For perspective, it is important to recognize
issues of scale. Some (but not all) genes within the decondensed euchromatic compartment will
be expressed, but this is a small fraction of the total open chromatin in this region, and pack-
aging changes for individual active genes occur at a much smaller scale than the formation of
the compartment. Increasing evidence supports that regional formation of heterochromatin is
not driven by the off state of individual genes. For example, during initiation of X inactivation
(induced by XIST RNA), the large, condensed Barr body forms before chromosome-wide gene
silencing (157), and Polycomb complexes (PRC1) can mediate long-range DNA interactions to
form heterochromatin compartments independent of local histone modifications and gene repres-
sion (18). Similarly, the initiation of the nuclear heterochromatic compartment forms in two- to
four-cell embryos before any cell type–specific changes in gene expression (25, 85).
These cytologically visible nuclear compartments are at a larger scale than structures detected
by Hi-C (chromosome conformation capture), which uses DNA cross-linking to investigate se-
quence organization in nuclei. This approach identifies topologically associating domains (TADs)
or sub-TADs and the larger A and B compartments (109). TADs are small intrachromosomal
self-interacting regions that are tethered by CTCF binding sites. Notably, more than 95% of
mammalian CTCF sites are derived from TEs (SINEs, LINEs, and long terminal repeats), and
almost all disease-associated STRs localize with CTCF boundaries (79). Although recent studies
have found increasing complexity to Hi-C structures, they are not clearly linked to euchro-
matin/heterochromatin packaging. However, TADs are bundled into A and B compartments
(of variable size, ∼1 Mb or more), which correspond to euchromatin (A) and heterochromatin
(B1/B2) bundles. While each A and B compartment reflects packaging well above the gene level,
the cytological-scale nuclear compartments are built by congregation of numerous A-with-A and
B-with-B Hi-C compartments.
We hypothesize below that the repetitive sequences that make up much of the fabric of a
chromosomal region are related to, and likely play a role in, forming heterochromatin versus
euchromatin regions. Before discussing how this relates to the karyotypic organization of gene
and repeat sequences, we summarize the important point that repeats are abundant not only in
DNA but also in nuclear RNA.

4.2. Abundant Repeat-Rich Scaffold RNA Physically Supports Open Chromatin


in Nuclear Territories
Recent evidence indicates that the unexplained length of pre-mRNA, lncRNA, and long inter-
genic ncRNA (lincRNA), much of which is repeat sequences, serves a structural role in nuclear
chromosome territories. Specifically, this RNA supports open euchromatin structure. C0 t-1 DNA
(the most highly repetitive genomic fraction; Figure 1a) is typically used as a cold competitor
to mask the “uninteresting” repeats. However, one study used labeled C0 t-1 DNA as a probe to
examine repeats in RNA by molecular cytology and made several unanticipated findings (73).
Since repeat-rich introns are generally rapidly degraded upon cotranscriptional splicing, the
bright, robust signal indicated a surprising abundance of repeats in RNA throughout the nucleus
(Figure 3e). C0 t-1 RNA is excluded from the peripheral heterochromatin and the inactive
X chromosome coated by XIST RNA. Analysis of a mouse/human hybrid cell with one human
[Link] • Repeat Genome Functions in Nuclear Structure 61
chromosome showed that the human C0 t-1 RNA remains tightly localized to the human chro-
mosome territory, unlike for excised introns, appearing remarkably similar to the XIST RNA
territory that covers the inactive X chromosome territory (Figure 3f ). Surprisingly, the localized
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

C0 t-1 RNA territory remained after prolonged transcriptional inhibition but could be rapidly dis-
persed by disrupting a nuclear scaffold protein, causing chromatin condensation (e.g., 98), which
suggested that it could be a euchromatic structural RNA. This idea was supported by several
studies showing that disruption of nuclear RNA causes cytological chromatin condensation and
implicating HnRNP-U (heterogeneous nuclear ribonucleoprotein U)/SAF-A (scaffold attach-
ment factor A) or similar proteins, which have both DNA- and RNA-binding domains, as being
involved (reviewed in 115).
To identify the RNA sequences involved in nuclear chromosome structure, a biochemical frac-
tionation procedure was developed to isolate nuclear scaffold RNAs that remain insoluble after
removal of histones and DNA (39) (Figure 3e). The procedure extracts most nuclear RNA and
leaves just 15% that cofractionates with known architectural RNAs, XIST RNA, and NEAT1
RNA [which forms the scaffold for nuclear paraspeckles (32)]. The insoluble RNAs that remained
with the nuclear scaffold are composed almost entirely of long, repeat-rich C0 t-1 RNAs (pre-
mRNAs, lncRNAs, and lincRNAs), and repeat RNA sequences are found almost entirely in the
nuclear scaffold fraction (Figure 3g). This C0 t-1 heterogeneous nuclear RNA is associated with
known nuclear scaffold/matrix RNA-binding proteins [matrin 3, NuMa (nuclear mitotic appara-
tus), and SAF-A] and forms an RNA-binding protein meshwork that promotes open chromatin.
A recent report also directly showed that matrin 3 binds repeat RNAs, particularly L1 anti-
sense RNA, and that the disruption of RNA binding (by a mutation that causes amyotrophic
lateral sclerosis) causes aberrant chromatin condensation (171). Thus, evidence supports that long,
repeat-rich “junk” RNA is integral to maintaining euchromatin structure and that the repeats
within this RNA may play a key role.
This evidence that long, repeat-rich heterogeneous nuclear RNAs function in maintaining
open chromatin suggests an unanticipated role for intron sequences in euchromatin structure
around active genes. Introns often allow for alternative splicing; however, this does not explain
their excessive length [some over 50 kb (142)] or why 80–90% of pre-mRNA sequence (and
many lncRNAs) is noncoding and replete with repeats. Repeat-rich intronic RNAs, lncRNAs, and
lincRNAs might help stabilize the epigenetic state of euchromatin and may also explain recent
findings that revealed exceptionally long-lived RNAs, including pre-mRNAs and lncRNAs, in nu-
clei of terminally differentiated mouse neurons (173); these RNAs appear to be highly analogous
to human C0 t-1 scaffold RNAs (for commentary, see 104).
Interspersed repeat RNAs may also play protective roles throughout the nuclear genome in
response to stress, adding to the evidence of a function for repeats in the stress response, as
established for HSat3 RNA (detailed in Section 2.4.1). For example, in response to stress, Alu ele-
ments are expressed from their own RNAPIII promoter, and the transcripts directly bind RNAPII,
repressing global transcription (116). The subsequent widespread loss of transcription would oth-
erwise cause deleterious chromatin condensation, but, surprisingly, new transcription of long,
intergenic, repeat-rich C0 t-1 RNAs is induced upon stress, including in response to osmotic shock
(159) or reversible transcriptional arrest (39). These extremely long intergenic transcripts have
been termed DOGS (downstream of genes) and suggested to play a role in protecting chromatin
from collapse by maintaining euchromatic C0 t-1 RNAs. This idea was supported by the observa-
tion that despite the arrest of genic transcription during the stress response, the new intergenic
transcription maintained C0 t-1 scaffold RNA levels (39), and this was required to avoid chromatin
collapse. These findings highlight that different components of the noncoding repeatome play a
role in response to stress.

62 Hall • Smith • Lawrence


Table 1 Characteristics of chromosome bands
G (Giemsa dark) R (reverse, or Giemsa light)
Sequence AT rich GC rich
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

CpG content CpG poor CpG rich


Gene density Gene rich Gene poor
Gene type Housekeeping Tissue specific
Replication timing Earlier Later
TE density LINE enriched SINE enriched

Abbreviations: LINE, long interspersed nuclear element; SINE, short interspersed nuclear element; TE, transposable
element.

4.3. Cytogenetic Bands Reveal High-Level Organization of Genes


and Interspersed Repeats
High-level organization of the genome sequence is apparent in the cytogenetic banding pattern
of mitotic chromosomes (Figure 1e), which reflects large blocks of chromatin (typically 5–10 Mb)
with distinct properties. Giemsa–trypsin staining produces a pattern of alternating dark G-bands
and light R-bands, which show certain differences in sequence content and biochemical properties
(15) (Table 1). Some evidence suggests that staining differences reflect different folding of DNA
loops into a proposed AT queue [along the chromosome axis (138)] and/or greater compaction of
G-band versus R-band chromosomal DNA (9, 68).
Important for our focus here is that the density and types of interspersed repeats, as well as
genes, also show a corresponding segmental distribution: SINE (Alu) elements are enriched in
R-bands, and LINE (L1) elements are enriched in G-bands (74, 100) (Figure 3h). There are
corresponding differences in CpG island distribution and late versus early replication between
chromosome bands (15) (Figure 3h; Table 1). Here, we discuss the segmental organization of
genes and repeats primarily in relation to G- and R-bands; however, we note that this binary
categorization is a simplification. As we briefly discuss below, some studies identify five categories
of bands based on the depth of Giemsa staining (64). Similarly, T-bands are a subtype of light
bands with an especially high density of genes and CpG islands (15). All of this reflects a segmental
organization of the genome sequence.
The functional significance for this cytological-scale organization has not been widely consid-
ered. Banding patterns are invariant between people (and cell types), but this does not necessarily
mean that this organization is unrelated to genome regulation. Relevant to this is that segmental
patterns of gene and repeat organization within synteny blocks are largely conserved (e.g., 38),
even though primary sequences of SINEs and LINEs are not. In addition, mobile SINEs do not
preferentially integrate into regions where they are most commonly found (90), and active L1s ap-
pear to prefer integration into R-bands (153) instead of G-bands, where L1 TEDS are enriched.
Thus, enrichment of LINEs in late-replicating segments and SINEs in gene-rich early-replicating
regions may be evolutionarily favored and conserved across species.
The essential co-opted functions of Alu elements with individual genes (e.g., TF binding, cov-
ered in Section 3) may partially explain why Alus are enriched in gene-rich bands. However, this
still raises the question of why genes—predominantly housekeeping genes—would be nonran-
domly clustered in large chromosome regions, corresponding to light R-bands. A strong rationale
for this clustering comes from the finding that genes and pre-mRNA metabolism are nonran-
domly organized within interphase nuclei. Within the euchromatin compartment, many active
genes preferentially distribute in ∼10–20 nuclear speckles (also known as SC35 domains) rich in
pre-mRNA metabolic and splicing factors (reviewed in 30, 75), and gene and Alu-rich R-bands

[Link] • Repeat Genome Functions in Nuclear Structure 63


are in spatial proximity to nuclear speckles, whereas adjacent dark bands are frequently condensed
at the periphery (143) (Figure 4a). The entire R-band in Figure 4a appears to be decondensed
(66), including regions not containing active genes, while other regions enriched for active genes
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

are expressed in close proximity to nuclear speckles (143). Gene clustering around nuclear hubs
that promote efficient gene expression further provides a functional rationale for the large-scale
clustered regional distribution of coding genes on chromosomes. This demonstrates a fundamen-
tal relationship between the segmental organization of the linear genome on chromosomes and
the structural organization of the functional genome in nuclei. It also explains why Alu-rich DNA
is more densely clustered around these same structures (29, 74) and why Alu-rich DNA is highly
correlated with regions of highest expression in the interphase nucleus (29).
Since Alu SINEs can contribute to the functions of individual genes, their enrichment in
gene-rich R-bands may simply reflect evolutionary conservation. However, their presence in gene-
rich regions would not require that Alu be strongly depleted from other regions (discussed in
Section 4.4), and the 1.1 million Alu TEDS that are more concentrated in the gene-rich segments
suggest greater Alu density than is easily explained by individual gene regulation. Hence, there
remains a question as to whether regional densities of Alu may be evolutionarily conserved for
additional, perhaps broader, contributions.

4.4. Highly Alu-Rich Segments Resist Condensation, and Alu Is Depleted


from L1 Heterochromatin
A priori, there is no clear reason why the density of Alu and L1 TEDS has an inverse distri-
bution in alternating multi-megabase chromosome segments, but this may well reflect distinct
contributions to nuclear compartmentalization in interphase. LINE repeats have long been sug-
gested to play a role in forming heterochromatin, including the spread of silencing on the
inactive X chromosome in female cells (31, 114). L1 is enriched in lamina-associated domains
(∼1–10-Mb DNA segments abutting the nuclear lamina and peripheral heterochromatin com-
partment), and increased L1 density is also seen in B compartments at the nuclear periphery
(reviewed in 106, 147) (Figure 4b,c). While heterochromatin is prevalent in the nuclear periphery
of most cell types, the overall nuclear pattern of heterochromatin is distinct and characteristic of
different cell types (Figure 3d). It has long been our view that these cytological patterns of chro-
matin architecture reflect the framework for coordinated genome-wide regulation of specific cell
types. The heterochromatin compartment will contain both constitutive heterochromatin, which
likely nucleates the compartment, and facultative heterochromatin, containing genes silenced in
specific cell types.
Since LINEs are prevalent throughout the genome, their modest enrichment in gene-poor G-
bands and heterochromatin could simply be because transposable LINEs were selected against in
gene-rich regions (R-bands). However, recent work has provided direct evidence for a functional
role of degenerate L1s in forming heterochromatin. L1 RNA, which is transiently expressed from
L1 TEDS in very early embryogenesis, is required for the de novo formation of the constitutive
heterochromatin compartment (before cell type–specific gene regulation begins) (85, 112, 113,
129). L1 RNA interacts with L1 DNA to help nucleate a compartment marked by specific chro-
matin modifications, and then the RNA must be silenced to maintain the heterochromatic state
(106, 149). Hence, L1 RNA induces changes that establish a stable heterochromatic state, similar
to XIST RNA triggering X chromosome heterochromatin. We note that this role of L1 RNA as
an inducer of heterochromatin formation is a distinct mechanism from C0 t-1 and L1 antisense
RNA physically supporting open euchromatin (Section 4.2).

64 Hall • Smith • Lawrence


a
Chromosome 17 R-band Chromosome 3 G-band Alu-rich
COL1A1 gene SC35 chromatin
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

SC35

L1-rich
chromatin

b A B
c
100 Tig-1
Percentage of repeat-rich

18%
80
compartments

60 77%

82%
40

20
23%
0
B1 rich L1 rich
L1 DNA Alu DNA L1 DNA Alu DNA
log2(B1/L1) >0 <0
Compartments called de novo
(p < 4 × 10−15)

d e 30
L1 DNA Alu DNA X chromosome
L1 LINEs (percentage of sequence)

Highest L1

25

Senescent cell Senescent SAHFs DAPI DNA


20 Chromosome 4

Chromosome 13
15
interactions
Long-range

Chromosome 20
0.8
Chromosome 21 Chromosome 17
0.4 Y chromosome Chromosome 16
10
0.0 Chromosome 19
Chromosome 22 Highest Alu

60 5
content (%) content (%)

Alu rich 0 5 10 15 20 25
40 L1 rich Alu SINEs (percentage of sequence)
Alu

20
0

60
40
L1

20
0
20 30 40
Chromosomal location (Mb)
(Caption appears on following page)

[Link] • Repeat Genome Functions in Nuclear Structure 65


Figure 4 (Figure appears on preceding page)
(a, left) DNA FISH image showing R-band DNA (17q21; green) with a COL1A1 gene (red) that closely associates with nuclear speckles
(stained for splicing factor SC35; blue). (Center) DNA FISH image showing that G-band DNA (red) is more condensed at the nuclear
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

periphery (SC35 speckles; green). (Right) Model showing that gene- and Alu-rich R-band DNA (light blue) is more intimately associated
with nuclear speckles than L1-rich, gene-poor G-band DNA (dark blue). Panel adapted with permission from Reference 143.
(b) Percentages of repeat-rich compartments as examined by Hi-C approaches. The most SINE-rich (B1 in mouse) compartments are
primarily A compartments (active), whereas the most LINE-rich (L1) are B compartments (inactive). Panel adapted from Reference
112 (CC BY 4.0). (c) DNA FISH images for L1 (red) and Alu (green) in human fibroblast nucleus, with the green signal outlined in the
second image to show Alu depletion at the periphery (DAPI; blue). Panel adapted with permission from Reference 74. (d, top) DNA
FISH images of a senescent fibroblast nucleus for L1 (red) and Alu (green) DNA. On the right, DAPI DNA shows dense
heterochromatin foci (SAHFs). (Bottom) Ideogram of 50 Mb of chromosome 4 aligned with the graph, showing changes in long-range
(>10 Mb) intrachromosomal Hi-C interactions between senescence and growing cells (log2 ). Each dot represents 100 kb and is colored
according to Giemsa-band designations; the black arrow indicates a condensing L1-rich R-band, and the blue arrow indicates an
R-band containing Alu-rich peaks that resist condensation. Also shown are the percentages of Alu and L1 content across the same
region, with the highest (90th percentile) Alu content outlined in green and the highest L1 content outlined in blue. Panel adapted with
permission from Reference 74. (e) Relative contributions (as percentages of total chromosome sequence) for L1 and Alu in all human
chromosomes. The insets show the X chromosome and chromosome 19 stained for L1 (red) and Alu (green) DNA. Panel adapted with
permission from Reference 74. Abbreviations: DAPI, 4′ ,6-diamidino-2-phenylindole; FISH, fluorescence in situ hybridization; L1, long
interspersed nuclear element 1; LINE, long interspersed nuclear element; SAHF, senescence-associated heterochromatic focus; SINE,
short interspersed nuclear element.

Recent evidence indicates that the depletion of Alu-rich peaks, not just L1 density, may be
an important factor influencing whether there is condensation of a region (74). In many primary
human fibroblasts, the peripheral heterochromatin appears to be more clearly delineated by Alu
depletion than by L1 enrichment (74) (Figure 4c). The condensed L1-rich Barr body, at the core of
the inactive X chromosome, also excludes Alu-rich DNA. In addition, when senescent cells com-
pletely reorganize peripheral heterochromatin into senescence-associated heterochromatic foci
(SAHFs), Alu-rich regions are excluded from these condensed bodies as well (74) (Figure 4d).
Hi-C data analysis revealed long contiguous Alu peaks as the most striking variance in repeat dis-
tribution, and the Alu-peak region consistently countered chromosome compaction (as indicated
by increased long-range intrachromosomal interactions). Figure 4d shows a quantification of L1
and Alu (in 100-kb bins) that demonstrates two additional points. First, L1 density on this chro-
mosome (chromosome 4) is similarly high across both R- and G-bands, which can show similar
structural changes (e.g., the black arrow indicates a condensing L1-rich R-band); in contrast, the
R-band containing Alu-rich peaks (blue arrow) resists this change (condensation). Second, results
also show that architectural interactions change in unison across whole large (∼5–15 Mb) chro-
mosome segments; for example, DNA throughout the whole darkest G-band shows increased
long-range interactions, suggesting that the band’s architecture changes as a single structural
unit.
Evidence indicates that constitutive heterochromatin forms in the darkest G-bands (G-positive
bands 75–100), which are the most L1 rich but also the lowest in Alu. Data from an earlier study of
five different band classes (64) indicated that SINEs (as a percentage of total sequence) essentially
double between the darkest and lightest bands (from 8.4% to 15.6%), whereas LINEs decrease
by ∼25% (from 25.1% to 19.2%). Marked differences in Alu and L1 enrichment are also seen
for certain whole chromosomes that similarly differ in their propensity to form heterochromatin.
The X chromosome has the highest L1 content (Figure 4e), although the L1 density is lower
in the pseudoautosomal region that escapes gene silencing (8). In marked contrast, chromosome
19 is strikingly Alu rich and low in L1, has the highest gene density, and consistently resides in
the euchromatic nuclear interior (78). Chromosome 19 is also unusual in that it does not form
a heterochromatic SAHF in senescent cells (74). Furthermore, chromosome 19 is an outlier in
that it encodes a concentration of more than 250 zinc finger regulatory proteins (44), many of

66 Hall • Smith • Lawrence


which are regulated by SVA elements. Based on the singular nature of the repeat content and
gene content of chromosome 19, we suggest that this whole small chromosome may be uniquely
maintained as constitutive euchromatin in different human cell types (and potentially in higher
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

primates with conserved synteny).

4.5. Future Directions: The Language of Chromosome Bands and Small


Common Words in Developmental Regulation
In our view, we have only seen the tip of the iceberg when it comes to the meaningful biology
in the sea of repeat-rich “junk” that makes up so much of the human and many higher genomes.
Clearly specific repeats have often been co-opted in ways that influence the structure or function
of individual genes. Granted, repeat content does not simply correlate with organismal complex-
ity, and the abundant TE-derived repeats could be mostly evolutionary detritus. However, if we
assume this, and repeats en masse are generally screened from studies, then we could overlook
a potentially fundamental aspect of genome and developmental biology. Evidence cited above
provides precedent that certain repeat families are expressed or otherwise play a role in specific
contexts, such as in response to stress, during developmental changes, or specific diseases. Hence,
if we do not look in different cellular contexts for changing expression, structural interactions, or
epigenetic modifications across the repeatome, we will not find them.
There are many next questions, and we will end by highlighting two with relevance to the
language of chromosome bands. As noted above, there are different flavors of bands, with differ-
ent depths of staining, which likely reflect sequence content. Learning to decode the language of
segmental chromosome organization may well require distinct approaches that seek to identify
patterns of organization across larger regions than typically studied. This may also prove impor-
tant in genome regulation during development. As illustrated in Figure 3d, the cell type–specific
patterns of cytological genome organization likely reflect changes in the heterochromatic ver-
sus euchromatic state of certain facultative regions. Changes in the expression of specific genes
can certainly occur within the euchromatin compartment, regulated at the histone or nucleosome
level. However, since many cell type–specific genes are in L1-rich dark bands, these are likely
facultative heterochromatin, and thus regions that switch compartmentalization in different cell
types (reviewed in 106), which may relate to the sequence differences reflected in distinct subtypes
of chromosome bands. We suggest that many cell type–specific genes will be regulated within re-
gions that can change chromatin state, and not only by mechanisms of individual gene expression.
This concept has fundamental importance for genome regulation that merits more investigation.
There has been more research into Alu and L1 TEDS, so these have been our focus here.
However, this is a simplistic view, because there is more complexity of repeats in the noncoding
genome. Perhaps most overlooked and poorly studied are abundant small “common words”—
interspersed simple sequence repeats (i-SSRs). In addition to TF binding sites, SSRs can form
unusual DNA structures that could influence chromatin folding (81). Hence, we make a distinc-
tion between locus-specific small tandem repeats (e.g., triplet repeats, discussed in Section 2.1)
because common i-SSRs are interwoven throughout the genomic fabric and thus could influence
its regional packaging. Comprising ∼3% of the human genome (101), some i-SSRs are remark-
ably prevalent, for unknown reasons. Most notably, the 9-mer word ATATATATA occurs 100 times
more frequently than the median 9-mer word (119), which could be related to an AT queue in the
chromosome axis (138). Chromosomal distributions of i-SSRs have not been well-characterized;
however, there is evidence that SSR enrichment in a region may correlate with its regulation.
For example, a count of all 9-mer words in the genome revealed a striking 11-fold enrichment of
GATAGATAG that was interspersed across the 10-Mb region of the X chromosome that escapes

[Link] • Repeat Genome Functions in Nuclear Structure 67


X inactivation (119). While the twofold-lower density of L1 TEDS in this region is often cited
as evidence for L1 function in silencing, the SSR content is largely overlooked. In this review, we
have not focused on this less studied feature of the genome, but i-SSRs could prove important to
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

understanding the language of the genome in chromosome biology.

DISCLOSURE STATEMENT
The authors are not aware of any affiliations, memberships, funding, or financial holdings that
might be perceived as affecting the objectivity of this review.

ACKNOWLEDGMENTS
We appreciate the support of National Institutes of Health grant R35 GM122597 to J.B.L.

LITERATURE CITED
1. Almeida MV, Vernaz G, Putman ALK, Miska EA. 2022. Taming transposable elements in vertebrates:
from epigenetic silencing to domestication. Trends Genet. 38:529–53
2. Altemose N. 2022. A classical revival: human satellite DNAs enter the genomics era. Semin. Cell Dev.
Biol. 128:2–14
3. Altemose N, Logsdon GA, Bzikadze AV, Sidhwani P, Langley SA, et al. 2022. Complete genomic and
epigenetic maps of human centromeres. Science 376:eabl4178
4. Andrews G, Fan K, Pratt HE, Phalke N, Zoonomia Consort., et al. 2023. Mammalian evolution of human
cis-regulatory elements and transcription factor binding sites. Science 380:eabn7930
5. Ardeljan D, Steranka JP, Liu C, Li Z, Taylor MS, et al. 2020. Cell fitness screens reveal a conflict between
LINE-1 retrotransposition and DNA replication. Nat. Struct. Mol. Biol. 27:168–78
6. Armanios M. 2022. The role of telomeres in human disease. Annu. Rev. Genom. Hum. Genet. 23:363–81
7. Bagshaw ATM. 2017. Functional mechanisms of microsatellite DNA in eukaryotic genomes. Genome
Biol. Evol. 9:2428–43
8. Bailey JA, Carrel L, Chakravarti A, Eichler EE. 2000. Molecular evidence for a relationship between
LINE-1 elements and X chromosome inactivation: the Lyon repeat hypothesis. PNAS 97:6634–39
9. Bak AL, Jorgensen AL, Zeuthen J. 1981. Chromosome banding and compaction. Hum. Genet. 57:199–
202
10. Batzer MA, Deininger PL. 2002. Alu repeats and human genomic diversity. Nat. Rev. Genet. 3:370–79
11. Beck CR, Garcia-Perez JL, Badge RM, Moran JV. 2011. LINE-1 elements in structural variation and
disease. Annu. Rev. Genom. Hum. Genet. 12:187–215
12. Belancio VP, Hedges DJ, Deininger P. 2006. LINE-1 RNA splicing and influences on mammalian gene
expression. Nucleic Acids Res. 34:1512–21
13. Bersani F, Lee E, Kharchenko PV, Xu AW, Liu M, et al. 2015. Pericentromeric satellite repeat expansions
through RNA-derived DNA intermediates in cancer. PNAS 112:15148–53
14. Betancourt AJ, Wei KH, Huang Y, Lee YCG. 2024. Causes and consequences of varying transposable
element activity: an evolutionary perspective. Annu. Rev. Genom. Hum. Genet. 25:1–25
15. Bickmore WA. 2019. Patterns in the genome. Heredity 123:50–57
16. Bizzotto S. 2023. The human brain through the lens of somatic mosaicism. Front. Neurosci. 17:1172469
17. Bourque G, Burns KH, Gehring M, Gorbunova V, Seluanov A, et al. 2018. Ten things you should know
about transposable elements. Genome Biol. 19:199
18. Boyle S, Flyamer IM, Williamson I, Sengupta D, Bickmore WA, Illingworth RS. 2020. A central role
for canonical PRC1 in shaping the 3D nuclear landscape. Genes Dev. 34:931–49
19. Britten RJ, Davidson EH. 1969. Gene regulation for higher cells: a theory. Science 165:349–57
20. Britten RJ, Kohne DE. 1968. Repeated sequences in DNA: Hundreds of thousands of copies of DNA
sequences have been incorporated into the genomes of higher organisms. Science 161:529–40
21. Brockdorff N. 2018. Local tandem repeat expansion in Xist RNA as a model for the functionalisation of
ncRNA. Noncoding RNA 4:28

68 Hall • Smith • Lawrence


22. Brouha B, Schustak J, Badge RM, Lutz-Prigge S, Farley AH, et al. 2003. Hot L1s account for the bulk
of retrotransposition in the human population. PNAS 100:5280–85
23. Burn A, Roy F, Freeman M, Coffin JM. 2022. Widespread expression of the ancient HERV-K (HML-2)
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

provirus group in normal human tissues. PLOS Biol. 20:e3001826


24. Burns KH. 2017. Transposable elements in cancer. Nat. Rev. Cancer 17:415–24
25. Burton A, Brochard V, Galan C, Ruiz-Morales ER, Rovira Q, et al. 2020. Heterochromatin establish-
ment during early mammalian development is regulated by pericentromeric RNA and characterized by
non-repressive H3K9me3. Nat. Cell Biol. 22:767–78
26. Carone DM, Lawrence JB. 2013. Heterochromatin instability in cancer: from the Barr body to satellites
and the nuclear periphery. Semin. Cancer Biol. 23:99–108
27. Cervantes-Ayalc A, Ruiz Esparza-Garrido R, Velazquez-Flores MA. 2020. Long Interspersed Nuclear
Elements 1 (LINE1): the chimeric transcript L1-MET and its involvement in cancer. Cancer Genet.
241:1–11
28. Chakravarti D, LaBella KA, DePinho RA. 2021. Telomeres: history, health, and hallmarks of aging. Cell
184:306–22
29. Chang Y-C, Quinodoz SA, Brangwynne CP. 2024. Live imaging of Alu elements reveals non-uniform
euchromatin dynamics coupled to transcription. eLife 13:RP97537
30. Chen Y, Belmont AS. 2019. Genome organization around nuclear speckles. Curr. Opin. Genet. Dev.
55:91–99
31. Chow JC, Ciaudo C, Fazzari MJ, Mise N, Servant N, et al. 2010. LINE-1 activity in facultative
heterochromatin formation during X chromosome inactivation. Cell 141:956–69
32. Clemson CM, Hutchinson JN, Sara SA, Ensminger AW, Fox AH, et al. 2009. An architectural role for
a nuclear noncoding RNA: NEAT1 RNA is essential for the structure of paraspeckles. Mol. Cell 33:717–
26
33. Cooper S, Dienstbier M, Hassan R, Schermelleh L, Sharif J, et al. 2014. Targeting Polycomb to peri-
centric heterochromatin in embryonic stem cells reveals a role for H2AK119u1 in PRC2 recruitment.
Cell Rep. 7:1456–70
34. Copley KE, Shorter J. 2023. Repetitive elements in aging and neurodegeneration. Trends Genet. 39:381–
400
35. Cordaux R, Batzer MA. 2009. The impact of retrotransposons on human genome evolution. Nat. Rev.
Genet. 10:691–703
36. Corless S, Hocker S, Erhardt S. 2020. Centromeric RNA and its function at and beyond centromeric
chromatin. J. Mol. Biol. 432:4257–69
37. Corneo G, Ginelli E, Polli E. 1967. A satellite DNA isolated from human tissues. J. Mol. Biol. 23:619–
22
38. Cournac A, Koszul R, Mozziconacci J. 2016. The 3D folding of metazoan genomes correlates with the
association of similar repetitive elements. Nucleic Acids Res. 44:245–55
39. Creamer KM, Kolpa HJ, Lawrence JB. 2021. Nascent RNA scaffolds contribute to chromosome
territory architecture and counter chromatin compaction. Mol. Cell 81:3509–25.e5
40. Criscione SW, Theodosakis N, Micevic G, Cornish TC, Burns KH, et al. 2016. Genome-wide
characterization of human L1 antisense promoter-driven transcripts. BMC Genom. 17:463
41. Davidson EH, Britten RJ. 1973. Organization, transcription, and regulation in the animal genome.
Q. Rev. Biol. 48:565–613
42. De Cecco M, Ito T, Petrashen AP, Elias AE, Skvir NJ, et al. 2019. L1 drives IFN in senescent cells and
promotes age-associated inflammation. Nature 566:73–78
43. de Lange T. 2018. Shelterin-mediated telomere protection. Annu. Rev. Genet. 52:223–47
44. Dehal P, Predki P, Olsen AS, Kobayashi A, Folta P, et al. 2001. Human chromosome 19 and related
regions in mouse: conservative and lineage-specific evolution. Science 293:104–11
45. Deininger P. 2011. Alu elements: know the SINEs. Genome Biol. 12:236
46. Denli AM, Narvaiza I, Kerman BE, Pena M, Benner C, et al. 2015. Primate-specific ORF0 contributes
to retrotransposon-mediated diversity. Cell 163:583–93
47. Depienne C, Mandel JL. 2021. 30 years of repeat expansion disorders: What have we learned and what
are the remaining challenges? Am. J. Hum. Genet. 108:764–85

[Link] • Repeat Genome Functions in Nuclear Structure 69


48. Dervan E, Bhattacharyya DD, McAuliffe JD, Khan FH, Glynn SA. 2021. Ancient adversary—HERV-K
(HML-2) in cancer. Front. Oncol. 11:658489
49. Dewannieux M, Esnault C, Heidmann T. 2003. LINE-mediated retrotransposition of marked Alu
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

sequences. Nat. Genet. 35:41–48


50. Dixit M, Ansseau E, Tassin A, Winokur S, Shi R, et al. 2007. DUX4, a candidate gene of facioscapulo-
humeral muscular dystrophy, encodes a transcriptional activator of PITX1. PNAS 104:18157–62
51. Du AY, Chobirko JD, Zhuo X, Feschotte C, Wang T. 2024. Regulatory transposable elements in the
encyclopedia of DNA elements. Nat. Commun. 15:7594
52. Dumbovic G, Forcales SV, Perucho M. 2017. Emerging roles of macrosatellite repeats in genome
organization and disease development. Epigenetics 12:515–26
53. Ehrlich M. 2009. DNA hypomethylation in cancer cells. Epigenomics 1:239–59
54. Elisaphenko EA, Kolesnikov NN, Shevchenko AI, Rogozin IB, Nesterova TB, et al. 2008. A dual origin
of the Xist gene from a protein-coding gene and a set of transposable elements. PLOS ONE 3:e2521
55. Eppard M, Passos JF, Victorelli S. 2024. Telomeres, cellular senescence, and aging: past and future.
Biogerontology 25:329–39
56. Etchegaray E, Naville M, Volff JN, Haftek-Terreau Z. 2021. Transposable element-derived sequences
in vertebrate development. Mob. DNA 12:1
57. Fattash I, Rooke R, Wong A, Hui C, Luu T, et al. 2013. Miniature inverted-repeat transposable elements:
discovery, distribution, and activity. Genome 56:475–86
58. Feschotte C, Pritham EJ. 2007. DNA transposons and the evolution of eukaryotic genomes. Annu. Rev.
Genet. 41:331–68
59. Feusier J, Watkins WS, Thomas J, Farrell A, Witherspoon DJ, et al. 2019. Pedigree-based estimation of
human mobile element retrotransposition rates. Genome Res. 29:1567–77
60. Fioriniello S, Marano D, Fiorillo F, D’Esposito M, Della Ragione F. 2020. Epigenetic factors that control
pericentric heterochromatin organization in mammals. Genes 11:595
61. Fonseca-Carvalho M, Verissimo G, Lopes M, Ferreira D, Louzada S, Chaves R. 2024. Answering the
cell stress call: satellite non-coding transcription as a response mechanism. Biomolecules 14:124
62. Franklin JM, Dubocanin D, Chittenden C, Barillas A, Lee RJ, et al. 2024. Human Satellite 3 DNA
encodes megabase-scale transcription factor binding platforms. bioRxiv 2024.10.22.616524. [Link]
org/10.1101/2024.10.22.616524
63. Fueyo R, Judd J, Feschotte C, Wysocka J. 2022. Roles of transposable elements in the regulation of
mammalian transcription. Nat. Rev. Mol. Cell Biol. 23:481–97
64. Furey TS, Haussler D. 2003. Integration of the cytogenetic map with the draft human genome sequence.
Hum. Mol. Genet. 12:1037–44
65. Georgiou I, Noutsopoulos D, Dimitriadou E, Markopoulos G, Apergi A, et al. 2009. Retrotranspo-
son RNA expression and evidence for retrotransposition events in human oocytes. Hum. Mol. Genet.
18:1221–28
66. Gharesouran J, Hosseinzadeh H, Ghafouri-Fard S, Taheri M, Rezazadeh M. 2021. STRs: ancient
architectures of the genome beyond the sequence. J. Mol. Neurosci. 71:2441–55
67. Gianfrancesco O, Geary B, Savage AL, Billingsley KJ, Bubb VJ, Quinn JP. 2019. The role of SINE-
VNTR-Alu (SVA) retrotransposons in shaping the human genome. Int. J. Mol. Sci. 20:5977
68. Gilbert N, Boyle S, Fiegler H, Woodfine K, Carter NP, Bickmore WA. 2004. Chromatin architecture
of the human genome: Gene-rich domains are enriched in open chromatin fibers. Cell 118:555–66
69. Ginelli E, Corneo G. 1976. The organization of repeated DNA sequences in the human genome.
Chromosoma 56:55–68
70. Gjerstorff MF. 2020. Novel insights into epigenetic reprogramming and destabilization of pericen-
tromeric heterochromatin in cancer. Front. Oncol. 10:594163
71. Greco A, Goossens R, van Engelen B, van der Maarel SM. 2020. Consequences of epigenetic
derepression in facioscapulohumeral muscular dystrophy. Clin. Genet. 97:799–814
72. Hall LL, Byron M, Carone DM, Whitfield TW, Pouliot GP, et al. 2017. Demethylated HSATII
DNA and HSATII RNA foci sequester PRC1 and MeCP2 into cancer-specific nuclear bodies. Cell Rep.
18:2943–56

70 Hall • Smith • Lawrence


73. Hall LL, Carone DM, Gomez AV, Kolpa HJ, Byron M, et al. 2014. Stable C0 T-1 repeat RNA is abundant
and is associated with euchromatic interphase chromosomes. Cell 156:907–19
74. Hall LL, Creamer KM, Byron M, Lawrence JB. 2024. Cytogenetic bands and sharp peaks of Alu underlie
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

large-scale segmental regulation of nuclear genome architecture. Nucleus 15:2400525


75. Hall LL, Smith KP, Byron M, Lawrence JB. 2006. Molecular anatomy of a speckle. Anat. Rec. A 288:664–
75
76. Hancks DC, Kazazian HH Jr. 2010. SVA retrotransposons: evolution and genetic instability. Semin.
Cancer Biol. 20:234–45
77. Hannan AJ. 2018. Tandem repeats mediating genetic plasticity in health and disease. Nat. Rev. Genet.
19:286–98
78. Harris RA, Raveendran M, Worley KC, Rogers J. 2020. Unusual sequence characteristics of human
chromosome 19 are conserved across 11 nonhuman primates. BMC Evol. Biol. 20:33
79. Haws SA, Simandi Z, Barnett RJ, Phillips-Cremins JE. 2022. 3D genome, on repeat: higher-order
folding principles of the heterochromatinized repetitive genome. Cell 185:2690–707
80. Hayward A, Gilbert C. 2022. Transposable elements. Curr. Biol. 32:R904–9
81. Herbert A. 2020. Simple repeats as building blocks for genetic computers. Trends Genet. 36:739–50
82. Hohn O, Hanke K, Bannert N. 2013. HERV-K(HML-2), the best preserved family of HERVs:
endogenization, expression, and implications in health and disease. Front. Oncol. 3:246
83. Hoyt SJ, Storer JM, Hartley GA, Grady PGS, Gershman A, et al. 2022. From telomere to telomere: the
transcriptional and epigenetic state of human repeat elements. Science 376:eabk3112
84. Iwasaki Y, Wada K, Wada Y, Abe T, Ikemura T. 2013. Notable clustering of transcription-factor-binding
motifs in human pericentric regions and its biological significance. Chromosome Res. 21:461–74
85. Jachowicz JW, Bing X, Pontabry J, Boskovic A, Rando OJ, Torres-Padilla ME. 2017. LINE-1 activa-
tion after fertilization regulates global chromatin accessibility in the early mouse embryo. Nat. Genet.
49:1502–10
86. Jjingo D, Conley AB, Wang J, Marino-Ramirez L, Lunyak VV, Jordan IK. 2014. Mammalian-wide in-
terspersed repeat (MIR)-derived enhancers and the regulation of human gene expression. Mob. DNA
5:14
87. Jjingo D, Huda A, Gundapuneni M, Marino-Ramirez L, Jordan IK. 2011. Effect of the transposable
element environment of human genes on gene length and expression. Genome Biol. Evol. 3:259–71
88. Johnson WE. 2019. Origins and evolutionary consequences of ancient endogenous retroviruses. Nat.
Rev. Microbiol. 17:355–70
89. Jones KW. 1970. Chromosomal and nuclear location of mouse satellite DNA in individual cells. Nature
225:912–15
90. Jurka J, Kohany O, Pavlicek A, Kapitonov VV, Jurka MV. 2004. Duplication, coclustering, and selection
of human Alu retrotransposons. PNAS 101:1268–72
91. Kale S, Boopathi R, Belotti E, Lone IN, Graies M, et al. 2023. The CENP-A nucleosome: where and
when it happens during the inner kinetochore’s assembly. Trends Biochem. Sci. 48:849–59
92. Kano H, Godoy I, Courtney C, Vetter MR, Gerton GL, et al. 2009. L1 retrotransposition occurs mainly
in embryogenesis and creates somatic mosaicism. Genes Dev. 23:1303–12
93. Kelley D, Rinn J. 2012. Transposable elements reveal a stem cell-specific class of long noncoding RNAs.
Genome Biol. 13:R107
94. Khan H, Smit A, Boissinot S. 2006. Molecular evolution and tempo of amplification of human LINE-1
retrotransposons since the origin of primates. Genome Res. 16:78–87
95. Kim S, Shin W, Lee YM, Mun S, Han K. 2020. Differential expressions of L1-chimeric transcripts in
normal and matched-cancer tissues. Anal. Biochem. 600:113769
96. Kishikawa T, Otsuka M, Yoshikawa T, Ohno M, Ijichi H, Koike K. 2016. Satellite RNAs promote
pancreatic oncogenic processes via the dysfunction of YBX1. Nat. Commun. 7:13006
97. Kit S. 1961. Equilibrium sedimentation in density gradients of DNA preparations from animal tissues.
J. Mol. Biol. 3:711–16
98. Kolpa HJ, Creamer KM, Hall LL, Lawrence JB. 2022. SAF-A mutants disrupt chromatin structure
through dominant negative effects on RNAs associated with chromatin. Mamm. Genome 33:366–81

[Link] • Repeat Genome Functions in Nuclear Structure 71


99. Konkel MK, Walker JA, Hotard AB, Ranck MC, Fontenot CC, et al. 2015. Sequence analysis and char-
acterization of active human Alu subfamilies based on the 1000 Genomes Pilot Project. Genome Biol.
Evol. 7:2608–22
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

100. Korenberg JR, Rykowski MC. 1988. Human genome organization: Alu, LINES, and the molecular
structure of metaphase chromosome bands. Cell 53:391–400
101. Lander ES, Linton LM, Birren B, Nusbaum C, Zody MC, et al. 2001. Initial sequencing and analysis of
the human genome. Nature 409:860–921
102. Landers CC, Rabeler CA, Ferrari EK, D’Alessandro LR, Kang DD, et al. 2021. Ectopic expression of
pericentric HSATII RNA results in nuclear RNA accumulation, MeCP2 recruitment, and cell division
defects. Chromosoma 130:75–90
103. Larsen PA, Hunnicutt KE, Larsen RJ, Yoder AD, Saunders AM. 2018. Warning SINEs: Alu elements,
evolution of the human brain, and the spectrum of neurological disease. Chromosome Res. 26:93–111
104. Lawrence J, Hall L. 2024. Exceptionally long-lived nuclear RNAs. Science 384:31–32
105. Lazaros L, Kitsou C, Kostoulas C, Bellou S, Hatzi E, et al. 2017. Retrotransposon expression
and incorporation of cloned human and mouse retroelements in human spermatozoa. Fertil. Steril.
107:821–30
106. Li S, Shen X. 2023. Long interspersed nuclear element 1 and B1/Alu repeats blueprint genome
compartmentalization. Curr. Opin. Genet. Dev. 80:102049
107. Li X, Bie L, Wang Y, Hong Y, Zhou Z, et al. 2024. LINE-1 transcription activates long-range gene
expression. Nat. Genet. 56:1494–502
108. Liang L, Cao C, Ji L, Cai Z, Wang D, et al. 2023. Complementary Alu sequences mediate enhancer-
promoter selectivity. Nature 619:868–75
109. Lieberman-Aiden E, van Berkum NL, Williams L, Imakaev M, Ragoczy T, et al. 2009. Comprehensive
mapping of long-range interactions reveals folding principles of the human genome. Science 326:289–93
110. Liu X, Wu B, Szary J, Kofoed EM, Schaufele F. 2007. Functional sequestration of transcription factor
activity by repetitive DNA. J. Biol. Chem. 282:20868–76
111. López-Martínez A, Soblechero-Martín P, de-la-Puente-Ovejero L, Nogales-Gadea G, Arechavala-
Gomeza V. 2020. An overview of alternative splicing defects implicated in myotonic dystrophy type I.
Genes 11:1109
112. Lu JY, Chang L, Li T, Wang T, Yin Y, et al. 2021. Homotypic clustering of L1 and B1/Alu repeats
compartmentalizes the 3D genome. Cell Res 31:613–30
113. Lu JY, Shao W, Chang L, Yin Y, Li T, et al. 2020. Genomic repeats categorize genes with distinct
functions for orchestrated regulation. Cell Rep. 30:3296–311.e5
114. Lyon MF. 2000. LINE-1 elements and X chromosome inactivation: a function for “junk” DNA? PNAS
97:6248–49
115. Marenda M, Lazarova E, Gilbert N. 2022. The role of SAF-A/hnRNP U in regulating chromatin
structure. Curr. Opin. Genet. Dev. 72:38–44
116. Mariner PD, Walters RD, Espinoza CA, Drullinger LF, Wagner SD, et al. 2008. Human Alu RNA is a
modular transacting repressor of mRNA transcription during heat shock. Mol. Cell 29:499–509
117. Marshall JN, Lopez AI, Pfaff AL, Koks S, Quinn JP, Bubb VJ. 2021. Variable number tandem repeats—
their emerging role in sickness and health. Exp. Biol. Med. 246:1368–76
118. McKerrow W, Wang X, Mendez-Dorantes C, Mita P, Cao S, et al. 2022. LINE-1 expression in cancer
correlates with p53 mutation, copy number alteration, and S phase checkpoint. PNAS 119:e2115999119
119. McNeil JA, Smith KP, Hall LL, Lawrence JB. 2006. Word frequency analysis reveals enrichment of
dinucleotide repeats on the human X chromosome and [GATA]n in the X escape region. Genome Res.
16:477–84
120. Miga KH. 2019. Centromeric satellite DNAs: hidden sequence variation in the human population. Genes
10:352
121. Miga KH, Alexandrov IA. 2021. Variation and evolution of human centromeres: a field guide and
perspective. Annu. Rev. Genet. 55:583–602
122. Miyata K, Imai Y, Hori S, Nishio M, Loo TM, et al. 2021. Pericentromeric noncoding RNA
changes DNA binding of CTCF and inflammatory gene expression in senescence and cancer. PNAS
118:e2025647118

72 Hall • Smith • Lawrence


123. Mocciaro E, Runfola V, Ghezzi P, Pannese M, Gabellini D. 2021. DUX4 role in normal physiology and
in FSHD muscular dystrophy. Cells 10:3322
124. Ninomiya K, Yamazaki T, Hirose T. 2023. Satellite RNAs: emerging players in subnuclear architecture
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

and gene regulation. EMBO J. 42:e114331


125. Nogalski MT, Shenk T. 2020. HSATII RNA is induced via a noncanonical ATM-regulated DNA damage
response pathway and promotes tumor cell proliferation and movement. PNAS 117:31891–901
126. Nogalski MT, Solovyov A, Kulkarni AS, Desai N, Oberstein A, et al. 2019. A tumor-specific endogenous
repetitive element is induced by herpesviruses. Nat. Commun. 10:90
127. Nurk S, Koren S, Rhie A, Rautiainen M, Bzikadze AV, et al. 2022. The complete sequence of a human
genome. Science 376:44–53
128. Pace JK II, Feschotte C. 2007. The evolutionary history of human DNA transposons: evidence for
intense activity in the primate lineage. Genome Res. 17:422–32
129. Pal M, Altamirano-Pacheco L, Schauer T, Torres-Padilla ME. 2023. Reorganization of lamina-
associated domains in early mouse embryos is regulated by RNA polymerase II activity. Genes Dev.
37:901–12
130. Payer LM, Burns KH. 2019. Transposable elements in human genetic disease. Nat. Rev. Genet. 20:760–
72
131. Payer LM, Steranka JP, Yang WR, Kryatova M, Medabalimi S, et al. 2017. Structural variants caused by
Alu insertions are associated with risks for many human diseases. PNAS 114:E3984–92
132. Pfaff AL, Singleton LM, Koks S. 2022. Mechanisms of disease-associated SINE-VNTR-Alus. Exp. Biol.
Med. 247:756–64
133. Podgornaya OI. 2022. Nuclear organization by satellite DNA, SAF-A/hnRNPU and matrix attachment
regions. Semin. Cell Dev. Biol. 128:61–68
134. Porter RL, Sun S, Flores MN, Berzolla E, You E, et al. 2022. Satellite repeat RNA expression in epithelial
ovarian cancer associates with a tumor-immunosuppressive phenotype. J. Clin. Investig. 132:e155931
135. Quinn JP, Bubb VJ. 2014. SVA retrotransposons as modulators of gene expression. Mob. Genet. Elem.
4:e32102
136. Raiz J, Damert A, Chira S, Held U, Klawitter S, et al. 2012. The non-autonomous retrotransposon SVA
is trans-mobilized by the human LINE-1 protein machinery. Nucleic Acids Res. 40:1666–83
137. Reis ALM, Rapadas M, Hammond JM, Gamaarachchi H, Stevanovski I, et al. 2023. The landscape of
genomic structural variation in Indigenous Australians. Nature 624:602–10
138. Saitoh Y, Laemmli UK. 1994. Metaphase chromosome structure: bands arise from a differential folding
path of the highly AT-rich scaffold. Cell 76:609–22
139. Senft AD, Macfarlan TS. 2021. Transposable elements shape the evolution of mammalian development.
Nat. Rev. Genet. 22:691–711
140. Sexton CE, Tillett RL, Han MV. 2022. The essential but enigmatic regulatory role of HERVH in
pluripotency. Trends Genet. 38:12–21
141. Shadle SC, Bennett SR, Wong CJ, Karreman NA, Campbell AE, et al. 2019. DUX4-induced bidirec-
tional HSATII satellite repeat transcripts form intranuclear double-stranded RNA foci in human cell
models of FSHD. Hum. Mol. Genet. 28:3997–4011
142. Shepard S, McCreary M, Fedorov A. 2009. The peculiarities of large intron splicing in animals. PLOS
ONE 4:e7853
143. Shopland LS, Johnson CV, Byron M, McNeil J, Lawrence JB. 2003. Clustering of multiple specific genes
and gene-rich R-bands around SC-35 domains: evidence for local euchromatic neighborhoods. J. Cell
Biol. 162:981–90
144. Smith KP, Byron M, Johnson C, Xing Y, Lawrence JB. 2007. Defining early steps in mRNA trans-
port: Mutant mRNA in myotonic dystrophy type I is blocked at entry into SC-35 domains. J. Cell Biol.
178:951–64
145. Smith KP, Hall LL, Lawrence JB. 2020. Nuclear hubs built on RNAs and clustered organization of the
genome. Curr. Opin. Cell Biol. 64:67–76
146. Smurova K, De Wulf P. 2018. Centromere and pericentromere transcription: roles and regulation. . .in
sickness and in health. Front. Genet. 9:674

[Link] • Repeat Genome Functions in Nuclear Structure 73


147. Solovei I, Thanisch K, Feodorova Y. 2016. How to rule the nucleus: divide et impera. Curr. Opin. Cell Biol.
40:47–59
148. Solovyov A, Vabret N, Arora KS, Snyder A, Funt SA, et al. 2018. Global cancer transcriptome quantifies
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

repeat element polarization between immunotherapy responsive and T cell suppressive classes. Cell Rep.
23:512–21
149. Stamidis N, Zylicz JJ. 2023. RNA-mediated heterochromatin formation at repetitive elements in
mammals. EMBO J. 42:e111717
150. Stein RA, DePaola RV. 2023. Human endogenous retroviruses: our genomic fossils and companions.
Physiol. Genom. 55:249–58
151. Stoyas CA, La Spada AR. 2018. The CAG-polyglutamine repeat diseases: a clinical, molecular, genetic,
and pathophysiologic nosology. Handb. Clin. Neurol. 147:143–70
152. Sullivan LL, Sullivan BA. 2020. Genomic and functional variation of human centromeres. Exp. Cell Res.
389:111896
153. Sultana T, van Essen D, Siol O, Bailly-Bechet M, Philippe C, et al. 2019. The landscape of L1 retrotrans-
posons in the human genome is shaped by pre-insertion sequence biases and post-insertion selection.
Mol. Cell 74:555–70.e7
154. Talley MJ, Longworth MS. 2024. Retrotransposons in embryogenesis and neurodevelopment. Biochem.
Soc. Trans. 52:1159–71
155. Taylor DJ, Eizenga JM, Li Q, Das A, Jenike KM, et al. 2024. Beyond the Human Genome Project: the
age of complete human genome sequences and pangenome references. Annu. Rev. Genom. Hum. Genet.
25:77–104
156. Timcheva K, Dufour S, Touat-Todeschini L, Burnard C, Carpentier MC, et al. 2022. Chromatin-
associated YTHDC1 coordinates heat-induced reprogramming of gene expression. Cell Rep. 41:111784
157. Valledor M, Byron M, Dumas B, Carone DM, Hall LL, Lawrence JB. 2023. Early chromosome
condensation by XIST builds A-repeat RNA density that facilitates gene silencing. Cell Rep. 42:112686
158. Vergnaud G, Denoeud F. 2000. Minisatellites: mutability and genome architecture. Genome Res. 10:899–
907
159. Vilborg A, Steitz JA. 2017. Readthrough transcription: How are DoGs made and what do they do? RNA
Biol. 14:632–36
160. Vukic M, Daxinger L. 2019. DNA methylation in disease: immunodeficiency, centromeric instability,
facial anomalies syndrome. Essays Biochem. 63:773–83
161. Vuoristo S, Bhagat S, Hyden-Granskog C, Yoshihara M, Gawriyski L, et al. 2022. DUX4 is a
multifunctional factor priming human embryonic genome activation. iScience 25:104137
162. Wada Y, Iwasaki Y, Abe T, Wada K, Tooyama I, Ikemura T. 2015. CG-containing oligonucleotides
and transcription factor-binding motifs are enriched in human pericentric regions. Genes Genet. Syst.
90:43–53
163. Walton EL, Francastel C, Velasco G. 2014. Dnmt3b prefers germ line genes and centromeric regions:
lessons from the ICF syndrome and cancer and implications for diseases. Biology 3:578–605
164. Wang J, Vicente-Garcia C, Seruggia D, Molto E, Fernandez-Minan A, et al. 2015. MIR retrotransposon
sequences provide insulators to the human genome. PNAS 112:E4428–37
165. Waring M, Britten RJ. 1966. Nucleotide sequence repetition: a rapidly reassociating fraction of mouse
DNA. Science 154:791–94
166. Wells JN, Feschotte C. 2020. A field guide to eukaryotic transposable elements. Annu. Rev. Genet. 54:539–
61
167. Wheeler TM, Thornton CA. 2007. Myotonic dystrophy: RNA-mediated muscle disease. Curr. Opin.
Neurol. 20:572–76
168. Xie Z, Liu C, Lu Y, Sun C, Liu Y, et al. 2022. Exonization of a deep intronic long interspersed nuclear
element in Becker muscular dystrophy. Front. Genet. 13:979732
169. Yamazaki T, Souquere S, Chujo T, Kobelke S, Chong YS, et al. 2018. Functional domains of NEAT1
architectural lncRNA induce paraspeckle assembly through phase separation. Mol. Cell 70:1038–53.e7
170. Zhang XO, Pratt H, Weng Z. 2021. Investigating the potential roles of SINEs in the human genome.
Annu. Rev. Genom. Hum. Genet. 22:199–218

74 Hall • Smith • Lawrence


171. Zhang Y, Cao X, Gao Z, Ma X, Wang Q, et al. 2023. MATR3-antisense LINE1 RNA meshwork scaffolds
higher-order chromatin organization. EMBO Rep. 24:e57550
172. Zhu J, Guo Q, Choi M, Liang Z, Yuen KWY. 2023. Centromeric and pericentric transcription and
Downloaded from [Link]. UGC-Infonet Digital Library Consortium (ar-367444) IP: [Link] On: Sun, 11 Jan 2026 15:19:53

transcripts: their intricate relationships, regulation, and functions. Chromosoma 132:211–30


173. Zocher S, McCloskey A, Karasinsky A, Schulte R, Friedrich U, et al. 2024. Lifelong persistence of nuclear
RNAs in the mouse brain. Science 384:53–59

[Link] • Repeat Genome Functions in Nuclear Structure 75

You might also like