0% found this document useful (0 votes)
6 views18 pages

Methuselah Proteins and Longevity Insights

This study analyzes 18 Methuselah (mth) protein variants from fruit flies, focusing on their evolutionary relationships, structural features, and roles in aging and longevity. The research identifies two major clades of mth proteins, reveals five functional subclasses, and highlights their similarities to G-protein-coupled receptors, suggesting their importance in signal transduction and cellular health. The findings emphasize the need for a quantitative approach to understanding Methuselah genes, contributing to insights on aging and longevity research.

Uploaded by

rameshdornala927
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views18 pages

Methuselah Proteins and Longevity Insights

This study analyzes 18 Methuselah (mth) protein variants from fruit flies, focusing on their evolutionary relationships, structural features, and roles in aging and longevity. The research identifies two major clades of mth proteins, reveals five functional subclasses, and highlights their similarities to G-protein-coupled receptors, suggesting their importance in signal transduction and cellular health. The findings emphasize the need for a quantitative approach to understanding Methuselah genes, contributing to insights on aging and longevity research.

Uploaded by

rameshdornala927
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

bioRxiv preprint doi: [Link] this version posted November 6, 2024.

The copyright holder for this


preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
Methuselah Proteins in Longevity: Unraveling Their Impact Through Mathematical
Genomics

Sk. Sarif Hassana,∗, Debaleena Nawnb,∗, Ankita Ghoshc , Moumita Silc , Arunava Goswamic , Pallab Basud , Kenneth
Lundstrome , Vladimir N. Uverskyf,∗∗
a Department of Mathematics, Pingla Thana Mahavidyalaya, Maligram, Paschim Medinipur, 721140, West Bengal, India
b Department of Computer Science and Engineering, Adamas University, Adamas Knowledge City, Barasat - Barrackpore
Road,, Jagannathpur, Kolkata, 700126, West Bengal, India
c Biological Science Division, Indian Statistical Institute, 203 Barrackpore Trunk Road, Kolkata 700108, West Bengal, India
d School of Physics, University of the Witwatersrand, Johannesburg, Braamfontein 2000, South Africa
e PanTherapeutics, Rte de Lavaux 49, CH1095 Lutry, Switzerland
f Department of Molecular Medicine, Morsani College of Medicine, University of South Florida, Tampa, FL 33612, USA

Abstract
This study provides a quantitative and comprehensive analysis of 18 Methuselah (mth) protein variants from fruit flies,
focusing on their evolutionary relationships, structural features, and functional roles in aging and longevity. Phylogenetic
analysis identified two major clades of mth proteins, with the first clade indicating conserved functions across Drosophila
species and the second clade reflecting gene duplication and diversification. The study found five distinct functional
subclasses of mth proteins through amino acid frequency and poly-string analyses, linked to their structural diversity
and role in longevity. Structural topology and post-translational modifications reveal similarities with G-protein-coupled
receptors (GPCRs), suggesting that mth proteins are crucial for signal transduction and cellular health. Variability
in propeptide cleavage sites and intrinsic protein disorder further highlight adaptive roles in signaling. The findings
underscore the importance of a quantitative and comprehensive approach to studying Methuselah genes, offering insights
into their functional versatility and evolutionary dynamics. This enhanced quantitative understanding contributes to
advancing research on aging and longevity.
Keywords: Methuselah (mth), Drosophila melanogaster, Aging, Functional diversity, G-protein-coupled receptors
(GPCRs), Intrinsic protein disorder.

1. Introduction

Aging is one of the most complex biological processes, influenced by both genetic and environmental factors [1,
2]. Single-gene mutations that extend lifespan are particularly valuable for understanding the molecular mechanisms
underlying aging and longevity determination [1, 3]. Gaining a detailed understanding of the molecular events involved
in aging will ultimately help reduce the impact of age-related diseases, improving human health and extending longevity
[4, 5].
The fruit fly Drosophila melanogaster serves as an excellent model system for dissecting the genetic and cellular basis of
crucial biological processes, such as aging [6]. This model allows researchers to uncover parallel mechanisms in vertebrates.
The conservation of human disease genes in Drosophila enables functional analysis of orthologs implicated in human aging
and age-related diseases [7]. Thus Drosophila melanogaster has proven to be one of the most valuable model systems for
studying the genetic determination of lifespan [8, 9]. Laboratory studies have demonstrated that its lifespan is highly
responsive to genetic manipulations such as induced mutations or artificial selection, with long-lived strains exhibiting
up to twice the lifespan of short-lived strains [10, 11]. Despite these findings, critical questions remain, such as whether
longevity is primarily governed by many genes of small effect or a few genes of large effect [12, 13, 14]. Furthermore, the
identification of aging genes through mutational analyses raises the question of whether these same genes contribute to
natural variation in lifespan [15]. Advances in genomic techniques are anticipated to enhance our understanding of the
complex genetic architecture of lifespan, particularly in comparing genes that extend lifespan in laboratory settings to
those affecting lifespan in natural populations [16].
Single-gene manipulations have led to the identification of candidate aging genes through extended longevity phe-
notypes, including the Insulin-like Receptor, chico, dFOXO, Indy, and methuselah in the model organism Drosophila
melanogaster [15, 17, 18]. Longevity and age-specific mortality patterns are complex traits that vary both within and
across species [19]. In model systems like Drosophila, several candidate aging genes have been uncovered through extended
longevity mutant phenotypes, such as the G-protein coupled receptor methuselah (mth) [17, 20, 21].

∗ These authors contributed equally in the work.


∗∗ Corresponding author: Vladimir N. Uversky
Email addresses: sksarifhassan@[Link] (Sk. Sarif Hassan), [Link]@[Link] (Debaleena Nawn),
ankitaghosh100@[Link] (Ankita Ghosh), moumitasil20@[Link] (Moumita Sil), srabanisopanarunava@[Link] (Arunava Goswami),
pallabbasu@[Link] (Pallab Basu), lundstromkenneth@[Link] (Kenneth Lundstrom), vuversky@[Link] (Vladimir N. Uversky )

Submitted to bioRxiv November 4, 2024


bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
Homozygous individuals with a p-element disruption at mth lived on average 35% longer than the parental strain and
showed notable resistance to oxidative stress, starvation, and heat stress [17, 22]. The mth gene encodes a G-protein
coupled receptor with seven hydrophobic regions indicative of transmembrane domains and an ectodomain with a ligand-
binding site [23]. Lifespan extension is also promoted by disrupting mth ligand activity [20, 24]. Mutations in the
stunted gene, which produces two mth peptide ligands, and constitutive expression of antagonist peptide ligands both
lead to increased longevity [20, 17, 25]. The mth gene is predicted to encode a protein homologous to several guanosine
triphosphate-binding protein-coupled seven-transmembrane domain receptors, suggesting that the organism may utilize
signal transduction pathways to modulate stress response and lifespan [26].
However, the mechanism by which mth regulates lifespan remains poorly understood [26]. To gain quantitative insight
into mth proteins, a study was conducted that offers a comprehensive analysis of 18 mth protein variants in fruit flies,
examining their evolutionary relationships, structural characteristics, and functional roles in relation to longevity.

2. Data acquisition

A list of 18 protein sequences of mth from various organisms of Fruit Fly were taken (Table 1). We utilized UniProt,
a comprehensive, high-quality, and freely accessible database of protein sequence and functional information, to extract
detailed information about each protein of interest [26].

Table 1: list of 18 mth proteins extracted from UNIPORT database


mth proteins from Fruit Fly

Sl. No UNIPROT ID Organism Number of Amino acids Functions


1 O97148 Drosophila melanogaster (Fruit fly) 514 Involved in biological aging and stress response. Essential for adult survival [26].
2 Q9VXD9 Drosophila melanogaster (Fruit fly) 676 Not known [27]
3 Q9GT50 Drosophila yakuba (Fruit fly) 517 Involved in biological aging and stress response. Essential for adult survival [28]
4 P83120 Drosophila simulans (Fruit fly) 515 Involved in biological aging and stress response. Essential for adult survival [29]
5 P83118 Drosophila melanogaster (Fruit fly) 532 Not known [27]
6 P83119 Drosophila melanogaster (Fruit fly) 488 Not known [27]
7 Q8SYV9 Drosophila melanogaster (Fruit fly) 533 Not known [27]
8 Q95NQ0 Drosophila simulans (Fruit fly) 536 Involved in biological aging and stress response. Essential for adult survival [30]
9 Q95NT6 Drosophila yakuba (Fruit fly) 536 Involved in biological aging and stress response. Essential for adult survival [30]
10 Q9V817 Drosophila melanogaster (Fruit fly) 517 Not known [27]
11 Q9V818 Drosophila melanogaster (Fruit fly) 511 Not known [27]
12 Q9VGG8 Drosophila melanogaster (Fruit fly) 496 Not known [27]
13 Q9VRN2 Drosophila melanogaster (Fruit fly) 518 Not known [27]
14 Q9VS77 Drosophila melanogaster (Fruit fly) 480 Not known [27]
15 Q9VSE7 Drosophila melanogaster (Fruit fly) 491 Not known [27]
16 Q9W0R5 Drosophila melanogaster (Fruit fly) 585 Not known [27]
17 Q9W0R6 Drosophila melanogaster (Fruit fly) 513 Not known [27]
18 Q9W0V7 Drosophila melanogaster (Fruit fly) 492 Not known [27]

3. Methods

Various quantitative features were extracted from 18 mth protein sequences. Methods of feature extractions were
discussed as follows.

3.1. Sequence homology and phylogeny


Multiple sequence alignment (MSA) was performed using Clustal Omega to identify conserved regions and assess
sequence similarity of 18 mth protein sequences[31, 32]. Consequently, a phylogenetic tree was constructed using the
Maximum Likelihood method in MEGA X, employing the JTT substitution model, chosen based on the Akaike Information
Criterion (AIC) [33]. The tree was visualized using iTOL, highlighting major evolutionary clades [31].

3.2. Amino acid frequency distribution across mth proteins


The frequency of amino acids in mth proteins were calculated using MATLAB in-build function and the total count
of each amino acid was then normalized by the total number of amino acids present in the dataset, providing a relative
frequency for each residue. This normalization was necessary to account for differences in sequence lengths [34, 35, 36].
The normalized frequency data were further analyzed to identify patterns and trends in amino acid usage in mth
proteins. Statistical measures, including mean, standard deviation, and variance, were calculated for each amino acid’s
frequency across the mth proteins. The frequency distribution was visualized using violin swarm-plots, highlighting any
amino acids with notably high or low representation. Correlation coefficient was also enumerated for all 18 mth proteins
based on frequency percentages of twenty amino acids.
The feature vectors of amino acid frequencies for the mth proteins were used to construct a distance matrix [36].
The Euclidean distance was calculated between each pair of feature vectors to quantify the differences in amino acid
composition. This distance matrix provided a measure of dissimilarity between the mth proteins based on their amino
acid frequency distributions. A phylogenetic tree was constructed using the distance matrix. The Neighbor-Joining (NJ)
method was selected due to its efficiency in handling large datasets and its capability to produce a topology that reflects
the input distance matrix accurately [37]. The NJ method was implemented using the MEGA X software, which facilitated
the tree construction and visualization [38].

2
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
3.3. Determining homogeneous poly-string frequency of amino acids in mth sequences
A homogeneous poly-string of length n is defined as a sequence consisting of n consecutive occurrences of a specific
amino acid, as described by Nawn et al. [36]. For instance, in the sequence ‘KKKLLKKLL’, there is one homogeneous
poly-string of K with a length of 3, another with a length of 2, and two homogeneous poly-strings of L, each with a length
of 2. It is crucial to note that only exclusive and exact occurrences of poly-strings of length n are counted.
To determine the maximum length of homogeneous poly-strings for all amino acids across all sequences, we recorded
the occurrences of homogeneous poly-strings for all possible lengths, from 1 up to the maximum length, for each amino
acid in a given protein sequence [36].

3.4. Determining structural topology and domains of methuselah proteins


A comprehensive detailed mapping of the structural topology, domains, and post-translational modifications from the
UniPort of each mth protein was extracted [39].

3.5. Prediction of proprotein convertase cleavage sites in mth proteins


Cleavage Site Prediction ProP - 1.0 webserver was used to predict the propeptide cleavage sites for arginine (R) and
lysine (K) in mth proteins [40]. Default configuration and relevant parameters were used and run the prediction by
inputting fasta mth sequences, which will output potential cleavage sites along with associated confidence scores.

3.6. Evaluating intrinsic protein disorder of mth protein sequences


Intrinsic protein disorder refers to the lack of a stable, well-defined three-dimensional structure in certain proteins
or protein regions under physiological conditions [41, 42, 43]. Proteins that exhibit this characteristic are known as in-
trinsically disordered proteins (IDPs). Unlike structured proteins, IDPs do not adopt a fixed conformation, remaining
flexible even under native conditions [44]. Despite their lack of a stable structure, IDPs are capable of performing diverse
biological functions [45].

To evaluate the per-residue disorder propensity in the mth protein sequences, we utilized PONDR® VSL2, a highly
accurate standalone disorder predictor [46]. This tool provides per-residue disorder predisposition scores ranging from 0
to 1, where a score of 0 signifies fully ordered residues and a score of 1 denotes fully disordered residues. Residues scoring
above the 0.5 threshold are considered disordered residues. Those with scores between 0.25 and 0.5 are classified as highly
flexible, and residues with scores between 0.1 and 0.25 are deemed moderately flexible [46].

3.7. Determining structural features of mth protein sequences


Protein structural features viz. solvent-accessible surface area (ASA), relative solvent accessibility (RSA), and dihedral
angles (ϕ and ψ), were predicted using NetSurfP-3.0 [47].

4. Results and analyses

4.1. Sequence homology and associated phylogeny of mth proteins


Sequence homology based phylogenetic relationship among 18 mth proteins reveals proteins Q95NQ0, Q9W0R5, and
Q95NT6 (Figure 1) are very much close to each other. This indicates that these proteins are likely orthologs, meaning
they are derived from a single ancestral gene present in the last common ancestor of Drosophila simulans i.e. DROSI,
Drosophila melanogaster i.e. DROME and Drosophila yakuba i.e. DROYA [48, 49]. The clades containing Drosophila
melanogaster proteins suggest significant gene duplication and diversification within this species, leading to a variety
of protein functions. The phylogenetic tree also indicates that O97148 (DROME) has much similarity with Q9GT50
(DROYA) and P83120 (DROSI). These relationships suggest shared functional and structural characteristics, highlighting
the importance of evolutionary conservation in understanding protein function.

3
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Figure 1: Phylogenetic relationship among mth sequences based on sequence homology

4.2. Amino acid frequency distribution across mth proteins


Normalized frequency of each amino acid was calculated across 18 mth proteins (Figure 2. The most abundant amino
acid is leucine (L), with an average frequency of 11.12%. High leucine content often correlates with structural roles due
to its hydrophobic nature, promoting tight packing in the protein core. Among all amino acids, percentage of leucine
is maximum in all mth sequences except Q9VSE7 where isoleucine (I) content is maximum (10.99%). The relatively
high presence of serine in most of the mth proteins (average frequency 7.27%) indicates potential roles in enzymatic
activity, possibly contributing to biochemical reactions necessary for cellular maintenance and longevity [50, 51]. Among
19 amino acids (excluding leucine), only serine (in Q8SYV9) and isoleucine (in P83119, Q9VS77, Q9VSE7, and Q9W0R6)
have normalized frequency greater than 9%. It was observed that hydrophobic amino acid content such as leucine (L),
isoleucine (I), valine (V), and phenylalanine (F) was comparatively high (average frequency 11.12%, 7.73%, 6.69%, and
6.59% respectively) and consequently, these proteins likely have a stable hydrophobic core, contributing to their overall
structural stability [52]. This stability is crucial for proteins involved in long-term cellular functions, possibly related to
the longevity associated with mth proteins [53]. Tryptophan (W) has the lowest average frequency at 2.18%. Despite
its low abundance, tryptophan is known for its role in protein-protein interactions and in stabilizing protein structures
through hydrophobic interactions [54].

Figure 2: Violin plot of the normalized frequency of each amino acid in 18 mth proteins

4
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
Based on the normalized frequency of amino acids in mth proteins, a phylogenetic relationship was drawn (Figure 3).
Similar to Figure 1, Q95NQ0 (DROSI), Q9W0R5 (DROME), and Q95NT6 (DROYA) are very close while Q9GT50
(DROYA) formed a cluster with O97148 (DROME) and P83120 (DROSI) based on normalized frequency of amino acids
also. However, Q9VGG8 is near to Q9V817 and Q9V818 in Figure 3, but this pattern is not found in Figure 1.

The clustering based on amino acid frequencies hints at structural conservation where certain amino acid compositions
are critical for the proteins’ roles in promoting longevity or resistance to stress [55]. This analysis provides a foundation
for further experimental validation and practical applications in understanding and leveraging the roles of these proteins
in promoting longevity and cellular maintenance [56].

Figure 3: Phylogenetic relationship among 18 mth proteins based on normalized (in %) amino acid frequency distributions

4.3. Homogeneous poly-string frequency of amino acids distribution across mth proteins
Table (2) provides the frequency of homogeneous poly-strings of different length l in 18 mth proteins. Maximum
length of homogeneous poly-strings turn out to be 8. Only Q9VXD9 has one poly-string of length 8 suggesting a unique
repetitive structure in this protein and this polystring is composed of Threonine (T). The threonine residues can serve
as sites for phosphorylation, impacting the protein’s activity and its involvement in signaling pathways related to aging
and stress response [57]. This polystring may also create a unique surface for binding interactions with other proteins
or ligands, facilitating the formation of multiprotein complexes crucial for neurogenesis [58]. Additionally, the flexibility
provided by the threonine-rich region allows Q9VXD9 to undergo conformational changes in response to environmental
cues, enhancing its regulatory capabilities [58]. However, this repetitive sequence could also predispose the protein to
misfolding or aggregation, potentially linking it to neurodegenerative processes. Overall, the threonine polystring is likely
integral to Q9VXD9’s roles in lifespan regulation and neuronal function [58].

No mth sequence has poly-string of length 6 or 7. Both of Q9GT50 and Q9V818 have single poly-string of length 5
consist of Leucine(L). Q9VXD9, Q9GT50 and Q8SYV9 exhibit single poly-string of length 4 consist of Serine (S), Leucine
and Arginine (R) respectively. No mth has poly-string of length 3 composed of Alanine (A), Cysteine (C), Aspartic acid
(D), Glycine (G), Histidine(H), Methionine(M), Arginine(R), and Tryptophan (W). All other amino acids have maximum
two poly-strings of length 3 in a single mth protein. NNN, PPP, TTT, and YYY (poly-string of Asparagine, Proline,
Threonine, and Tyrosine respectively having length 3) was present only in Q9VRN2, Q9VXD9, O97148, and Q9W0R6 re-
spectively. Among them, NNN, TTT, and YYY were found with single occurrence while two PPP was noticed in Q9VXD9.

High frequencies of single amino acids and the presence of longer homogeneous poly-strings in certain proteins suggest
specific structural motifs and functions that could be critical for the roles these proteins play in longevity [59, 60].

5
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Table 2: Frequency of homogeneous poly-strings of amino acids across mth proteins


mth Proteins l=1 l=2 l=3 l=4 l=5 l=8 mth Proteins l=1 l=2 l=3 l=4 l=5 l=8
O97148 450 26 4 0 0 0 Q9V817 452 31 1 0 0 0
Q9VXD9 540 50 8 1 0 1 Q9V818 446 27 2 0 1 0
P83120 456 25 3 0 0 0 Q9VGG8 410 40 2 0 0 0
Q9GT50 442 27 4 1 1 0 Q9VRN2 440 33 4 0 0 0
P83118 470 28 2 0 0 0 Q9VS77 427 25 1 0 0 0
P83119 438 22 2 0 0 0 Q9VSE7 441 22 2 0 0 0
Q8SYV9 444 41 1 1 0 0 Q9W0R5 515 29 4 0 0 0
Q95NQ0 474 25 4 0 0 0 Q9W0R6 442 28 5 0 0 0
Q95NT6 478 23 4 0 0 0 Q9W0V7 452 20 0 0 0 0

4.4. Structural topology and domains of methuselah proteins


It was noted that most mth proteins possess similar structural topology with signal peptides ranging from 17 to 37
amino acids and consistent chain lengths (Table 3). The presence of topological domains interspersed with transmem-
brane regions suggests that these proteins span the cell membrane multiple times, typical of G-protein-coupled receptors
(GPCRs) [61]. Furthermore, the number and position of transmembrane regions are quite consistent among the proteins,
typically showing 7 transmembrane regions. This is characteristic of GPCRs, which are involved in transmitting signals
from the outside to the inside of a cell. There is a significant overlap in the domains identified among the mth proteins,
indicating conserved regions essential for their function [62]. Several glycosylation sites are conserved among the proteins.
Glycosylation is crucial for protein folding, stability, and interactions, indicating that these modifications are important
for mth protein functionality [63, 64]. The presence of disulfide bonds across all proteins highlights the importance of
structural integrity and proper folding for their function [65].

The structural characteristics of mth proteins, resembling GPCRs, suggest they are involved in signal transduction path-
ways. These pathways are crucial for cellular responses to environmental stimuli, which can influence longevity [66].
Glycosylation and disulfide bonds are important for protein stability and function. Properly folded and stable proteins
are more likely to function effectively over time, potentially contributing to cellular health and longevity [67]. The con-
servation of domains across different mth proteins implies that they might have similar roles in the organism. This
functional redundancy could be a protective mechanism, ensuring that essential signaling pathways are maintained, which
could positively impact longevity [68]. GPCRs, including mth proteins, are often involved in stress response pathways.
Effective stress response mechanisms can help mitigate damage from environmental stressors, thereby promoting longevity.

The mth proteins exhibit highly conserved structural features and modifications that are typical of GPCRs [69, 70]. These
characteristics indicate their significant role in signal transduction pathways, which are crucial for maintaining cellular
homeostasis and responding to environmental changes. The integrity and functionality provided by glycosylation and
disulfide bonds further underscore their importance in cellular processes [71]. Given their roles in stress response and
signaling, mth proteins likely contribute to mechanisms that promote longevity, supporting the organism’s ability to cope
with environmental challenges and maintain cellular health over time [72].

6
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Table 3: Molecule Processing, Topology, Domains and Post-Translational Modifications (PTM) of mth proteins
Amino acid positions
Methuselah
Molecule processing Topology
Sl. No UNIPROT ID Signal peptide Chain Topological domain Transmembrane
1 O97148 1–24 25–514 25–218, 240–248, 270–278, 300–320, 342–370 219–239, 249–269, 279–299, 321–341
2 Q9VXD9 1–37 38–676 38–335, 357–365, 387–403, 425–441, 463–488, 510–540, 562, 584–676 336–356, 366–386, 404–424, 442–462, 489–509, 541–561, 563–583
3 Q9GT50 1–27 28–517 28–221, 243–251, 273–279, 301–323, 345–373, 395–427, 449–457, 479–517 222–242, 252–272, 280–300, 324–344, 374–394, 428–448, 458–478
4 P83120 1–24 25-515 25–218, 240–248, 270–278, 300–320, 342–370, 392–424, 446–454, 476–515 219–239, 249–269, 279–299, 321–341, 371-391, 425–445, 455–475
5 P83118 1–20 21–532 21–229, 251-262, 284–290, 312–339, 361–386, 408–439, 461–469, 491–532 230–250, 263-283, 291–311, 340–360, 387–407, 440–460, 470–490
6 P83119 1–17 18–488 18–215, 237–247, 269–283, 305–315, 337, 373, 395–416, 438–454, 476–488 216–236, 248–268, 284–304, 316–336, 374–394, 417–437, 455–475
7 Q8SYV9 1–23 24–533 24–242, 264–279, 301–303, 325–347, 369–395, 417–451, 473–480, 502–533 243–263, 280–300, 304–324, 348–368, 396–416, 452–472, 485–501
8 Q95NQ0 233–241, 263–273, 295–314, 336–365, 387–417, 439–449, 471–536 212–232, 242–262, 274–294, 315–335, 366–386, 418–438, 450-470
9 Q95NT6 1–210, 232–241, 263–273, 295–314, 336–365, 387–417, 439–449, 471–536 211–231, 242–262, 274–294, 315–335, 366-386, 418–438, 450–470
10 Q9V817 1–18 19–517 19–212, 234–242, 264–272, 294–319, 341–363, 385–414, 436–459, 481–517 213–233, 243–263, 273–293, 320–340, 364–384, 415–435, 460–480
11 Q9V818 1–21 22–511 22–215, 237–245, 267–273, 295–317, 339–366, 388-422, 444–452, 474–511 216–236, 246–266, 274–294, 318–338, 367–387, 423–443, 453–473
12 Q9VGG8 1–496 1–219, 241–246, 268–276, 298–327, 349–366, 388–411, 433–438, 460–496 220–240, 247–267, 277–297, 328–348, 367–387, 412–432, 439–459
13 Q9VRN2 1–26 27–518 27–220, 242–250, 272–279, 301–321, 343–371, 393–426, 448–455, 477–518 221–241, 251–271, 280–300, 322–342, 372–392, 427–447, 456–476
14 Q9VS77 1–20 21–480 21–202, 226–231, 255–263, 284–303, 327–356, 380–405, 429–437, 458–480 203–225, 232–254, 264–283, 304–326, 357–379, 406–428, 438–457
15 Q9VSE7 1–22 23–491 23–167, 189–222, 244–252, 274–325, 347–372, 394–434, 456–458, 480–491 168–188, 223–243, 253–273, 326–346, 373–393, 435–455, 459–479
16 Q9W0R5 1–32 33–585 33–250, 272–280, 302–312, 334–353, 375–404, 426–466, 48–498, 520–585 251–271, 281–301, 313–333, 354–374, 405–425, 467–487, 499–519
17 Q9W0R6 1–19 20–513 20–207, 229–242, 264–276, 298–314, 336–360, 382–403, 425–438, 460–513 208–228, 243–263, 277–297, 315–335, 361–381, 404–424, 439–459
18 Q9W0V7 1–21 22–492 22–218, 240–245, 267–282, 304–317, 339–362, 384–411, 433–441, 463–492 219–239, 246–266, 283–303, 318–338, 363–383, 412–432, 442–462
Methuselah Domains Post-Translational Modifications (PTM)
Sl. No UNIPROT ID InterPro Representative Domain Region Glycosylation Sites Disulfide bonds
1 O97148 29–209, 219–489 45, 109, 123, 170, 198 29–83, 85–90, 94–188, 95–106, 150–209
2 Q9VXD9 326–598 609–658 136, 165, 276, 486 198–283, 199–212, 198–283
3 Q9GT50 32–212, 215-492 48, 61, 126, 173, 201, 452 32–86, 88–93, 97–191, 98–109, 153–212
4 P83120 29–209, 219–489 45, 109, 123, 170, 198 29–83, 85–90, 94–188, 95–106, 150–209
5 P83118 26–204, 236–504 42, 110, 123, 166, 195, 227 26–80, 82–87, 91–184, 92–103, 145–204
6 P83119 27–209, 210–479 19, 34, 55, 141, 365 27–81, 83–88, 92–189. 93–104, 155–209
7 Q8SYV9 239–512 86–108 20, 30, 36, 47, 133, 178, 206 120–216
8 Q95NQ0 17–197, 203–484 487–506 24, 33, 103, 113, 118, 159, 184 17–71, 73–78, 82–177, 83–96, 138–197
9 Q95NT6 17–197, 204–484 487–506 24, 33, 103, 113, 118, 159, 184, 203 17–71, 73–78, 82–177, 83–96, 138–197
10 Q9V817 23–201, 205–477 39, 117, 165, 456 23–77, 79–84, 88–183, 89–100, 145–201
11 Q9V818 26–204, 208–486 42, 120, 168 26–80, 82–87, 91–186, 92–103, 148–204
12 Q9VGG8 215–473 82
13 Q9VRN2 31–211, 199–475 47, 111, 125, 201
14 Q9VS77 25–197, 211–471 40, 160, 170 25–78, 80–85, 89–179, 90–101
15 Q9VSE7 27–214, 218–482 18, 42, 248 27–80, 82–87, 92–103
16 Q9W0R5 56–236, 242–533 63, 72, 142, 152, 157, 198, 223 56–110, 112–117, 121–216, 122–135, 177–236
17 Q9W0R6 29–199, 232–473 36, 106, 125, 165 29–82, 84–89, 93–181, 94–107
18 Q9W0V7 30–202 37, 51, 129, 169, 192, 275, 360, 438 30–82, 84–89, 93–184, 94–107

4.5. Predicting arginine and lysine propeptide cleavage sites in mth proteins
Most proteins have a predicted signal peptide cleavage site, suggesting they are likely to be processed and have their
signal peptide removed, which is a typical post-transnational modification for secreted and membrane proteins [73]. On
the other side it was noticed that two proteins, Q95NQ0 and Q95NT6, do not have predicted signal peptide cleavage sites
(‘none’), indicating these proteins may not follow the typical secretory pathway or may lack a signal peptide altogether
(Table 4) [74]. The position of the cleavage site varies among proteins, with some occurring early (e.g., P83119 at position
17 and 18) and others much later (e.g., Q9VGG8 at position 43 and 44). The sequences around the cleavage sites also
vary, indicating diversity in the signal peptides and possibly in the mechanisms by which these proteins are processed
[75].
The majority of the mth proteins listed have “Arg(R)/Lys(K): 0,” indicating no predicted pro-peptide cleavage sites
for these amino acids [76]. This suggests that these proteins do not undergo cleavage at these sites during maturation, or
such cleavage is not a feature of their post-transnational modification [77]. Only two proteins, P83119 and Q9VXD9, have
a value of ‘1’ in the pro-peptide cleavage sites, indicating one predicted cleavage site. This suggests a specific site where
enzymatic cleavage could occur during the protein maturation process, possibly affecting protein activation, function, or
stability [78].

A prevalent feature among these mth proteins is the presence of a predicted signal peptide cleavage site, which is a
hallmark of proteins destined for secretion or localization to specific cellular compartments [79]. The lack of pro-peptide
cleavage sites for most proteins suggests that pro-peptide cleavage may not be a common or necessary step in their matu-
ration [80]. Furthermore, the variability in cleavage site positions and sequences indicates functional diversity among these
proteins, potentially reflecting different roles or mechanisms of action [81]. The few proteins with predicted pro-peptide
cleavage sites may have unique regulatory mechanisms, involving the removal of specific peptide segments to achieve
functional maturation [82].

The presence of signal peptide cleavage sites suggests these proteins are likely to be secreted or associated with
membranes, playing roles in inter cellular signaling, receptor functions, or other membrane-associated processes [83]. The
absence or presence of pro-peptide cleavage sites provides additional insight into the processing and functional regulation
of these proteins.

7
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Table 4: Predictions of arginine and lysine pro-peptide cleavage sites in mth proteins
Methuselah proteins Signal peptide cleavage site predicted (between pos) Pro-peptide cleavage sites predicted
O97148 24 and 25: SYA-DI Arg(R)/Lys(K): 0
P83118 20 and 21: VRS-RD Arg(R)/Lys(K): 0
P83119 17 and 18: TIA-KN Arg(R)/Lys(K): 1
P83120 24 and 25: SYA-DI Arg(R)/Lys(K): 0
Q8SYV9 23 and 24: ASA-QI Arg(R)/Lys(K): 0
Q95NQ0 none Arg(R)/Lys(K): 0
Q95NT6 none Arg(R)/Lys(K): 0
Q9GT50 27 and 28: TNA-AI Arg(R)/Lys(K): 0
Q9V817 18 and 19: SNA-EI Arg(R)/Lys(K): 0
Q9V818 21 and 22: SNA-EI Arg(R)/Lys(K): 0
Q9VGG8 43 and 44: VTS-HV Arg(R)/Lys(K): 0
Q9VRN2 26 and 27: SSA-EI Arg(R)/Lys(K): 0
Q9VS77 20 and 21: SEA-VI Arg(R)/Lys(K): 0
Q9VSE7 22 and 23: SNA-DI Arg(R)/Lys(K): 0
Q9VXD9 37 and 38: SLA-IE Arg(R)/Lys(K): 1
Q9W0R5 32 and 33: IPG-IP Arg(R)/Lys(K): 0
Q9W0R6 19 and 20: AKS-VE Arg(R)/Lys(K): 0
Q9W0V7 21 and 22: SWG-FH Arg(R)/Lys(K): 0

4.6. Intrinsic protein disorder of mth proteins


Per residue intrinsic protein disorder of mth proteins were calculated and presented in Figure 4. Intrinsic protein
disorder prediction analysis revealed a high variation in the intrinsic disorder in the C- and N-terminals of mth proteins
(Figure 4). It also noticed that the highest percentage of moderately flexible residues were present in all mth proteins
other than Q9V817, Q8SYV9, Q9V818, and P83118 in whcih highly flexible residues were dominating residues (Figure
5).

Proteins like Q9VXD9 (32.988%) has significant disordered regions, indicating potential roles in regulation and inter-
actions. Disordered regions are typically involved in signaling and can bind to multiple partners, enhancing functional
versatility. P83119 (28.074%) and Q9V817 (26.692%) show high flexibility, suggesting adaptability in their interactions
and potential for signaling roles. Flexibility can facilitate binding to diverse molecules, which is crucial for proteins in-
volved in multiple pathways. Furthermore, moderately flexible regions were consistent across proteins, around 20-30%,
balancing stability and flexibility, essential for proper function. Almost all mth proteins, except Q9VXD9 and P83119
possessed well-structured regions (‘others’ as mentioned in Figure 5).

The combination of disordered, flexible, and well-structured regions suggests that these proteins are well-adapted to
handle various cellular functions [84]. This adaptability is likely important for responding to stress and maintaining
cellular health, contributing to longevity [85]. It was noted that the presence of significant disordered and flexible regions
suggests these mth proteins can interact dynamically with other molecules, crucial for signaling and regulatory processes
[86]. Effective cellular signaling and regulation are vital for longevity, as they help maintain homeostasis and prevent
cellular damage [87].

Figure 4: Intrinsic protein disorder residue plots of mth sequences

8
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Figure 5: Percentage of intrinsically protein disorder regions of mth sequences

4.7. Methuselah protein structural features


For each amino acid in each mth protein, three-state secondary structure viz. relative solvent-accessible area (RSA),
and the dihedral angles ϕ and ψ were computed using NetSurfP-3.0. A snapshot of the predicted structure with RSA and
the dihedral angles ϕ and ψ of the reference mth sequence O97148 was presented (Figure 6).

Figure 6: Predicted secondary structural features with RSA and the dihedral angles ϕ and ψ for the mth protein O97148

9
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Figure 7: Structural features with RSA and the dihedral angles ϕ and ψ for all the mth protein sequences

10
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
Distribution of RSA, ϕ and ψ were presented in form of swarm-plots (Figures 8–10).

Figure 8: Distribution of RSA

Figure 9: Distribution of ϕ

11
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Figure 10: Distribution of ψ

Furthermore, mean and standard deviation (SD) of the three features viz. RSA, ϕ, and ψ for all mth proteins across all
amino acid residues were enumerated (Table 5).

Table 5: Mean and SD of RSA and ϕ and ψ for each mth protein
RSA Phi psi RSA Phi psi
Uniprot ID Uniprot ID
Mean SD Mean SD Mean SD Mean SD Mean SD Mean SD
O97148 0.32 0.23 -76.91 29.56 26.05 74.03 Q9V818 0.31 0.23 -74.24 33.32 23.63 75.15
P83118 0.31 0.22 -77.03 32.79 31.37 77.44 Q9VGG8 0.32 0.23 -73.26 30.93 19.95 71.52
P83119 0.30 0.21 -76.82 29.15 24.91 74.02 Q9VRN2 0.32 0.23 -75.58 30.28 20.63 73.05
P83120 0.32 0.23 -77.07 29.75 25.92 74.02 Q9VS77 0.30 0.22 -74.91 32.10 27.77 76.87
Q8SYV9 0.36 0.23 -74.96 29.49 23.65 70.31 Q9VSE7 0.28 0.21 -76.00 32.49 24.84 74.67
Q95NQ0 0.31 0.22 -76.54 30.70 24.27 75.15 Q9VXD9 0.37 0.23 -74.34 36.02 36.69 76.11
Q95NT6 0.31 0.22 -76.41 30.77 24.08 75.10 Q9W0R5 0.32 0.22 -75.61 34.34 26.54 75.74
Q9GT50 0.32 0.23 -76.34 31.88 27.40 75.80 Q9W0R6 0.32 0.21 -76.20 31.49 29.51 75.77
Q9V817 0.32 0.23 -74.79 31.64 22.75 74.49 Q9W0V7 0.31 0.22 -74.38 37.97 28.28 79.34

Based on this mean and standard deviation of the three features for all 18 mth proteins, a phylogenetic relationship was
drawn in Figure 11. From this dendrogram five clusters were developed and they are . . .

Cluster 1: Q9GT50, O97148, P83119, Q9VSE7, P83120, P83118.


Cluster 2: Q95NQ0, Q95NT6, Q9VS77, Q9W0R5, Q8SYV9.
Cluster 3: Q9VGG8, Q9VRN2, Q9V818, Q9W0V7.
Cluster 4: Q9VXD9, Q9W0R6.
Cluster 5: Q9V817.

12
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.

Figure 11: Phylogenetic relationship among mth proteins based on RSA, ϕ, and ψ.

The mean RSA values range from approximately 0.28 to 0.37. This suggests that the majority of these proteins have
moderate solvent accessibility, indicating that they are neither completely buried nor fully exposed. Mean RSA values of
the most of the mth proteins hover around 0.31 to 0.32, reflecting a general tendency for these proteins to have similar
levels of exposure to the solvent, which could be important for their interactions with other molecules [88]. Furthermore,
the moderate RSA means indicate a balance between hydrophobic and hydrophilic regions, which is crucial for protein
folding, stability, and function. Proteins with such accessibility are likely to engage in diverse interactions, both hydropho-
bic and hydrophilic [88, 89]. The distribution of RSA among the 18 proteins indicates a general trend towards moderate
accessibility with low variability, suggesting a common structural feature that may be critical for their biological functions
[90].

5. Discussion and Concluding Remarks

This study provides a quantitative and comprehensive analysis of 18 Methuselah (mth) protein variants from fruit flies,
focusing on their evolutionary relationships, structural features, and functional roles related to longevity.

Phylogenetic analysis reveals two major clades among the mth proteins. The first major clade, including Q9W0R5,
Q95NT6, and Q95NQ0, suggests these proteins are orthologs derived from a common ancestral gene, indicating conserved
functions across species such as Drosophila melanogaster, Drosophila yakuba, and Drosophila simulans [91]. The second
major clade reflects more complex evolutionary dynamics, likely due to gene duplication and subsequent functional di-
vergence. Within Drosophila melanogaster, multiple sub-clades suggest extensive gene duplication and diversification,
potentially leading to diverse protein functions [92]. The close relationships between proteins such as O97148, Q9GT50,
and P36120 highlight conserved functional and structural characteristics [93].

Amino acid frequency analysis identified five distinct clusters of mth proteins, suggesting functional subclasses with
varying structural features essential for longevity [94]. Poly-string frequency analysis further emphasizes the structural
diversity and functional roles of mth proteins, with notable variations in single amino acid frequencies and repetitive
motifs indicating specific functional adaptations [93].

Structural topology and post-translational modifications reveal that mth proteins share a common topology with signal
peptides and multiple transmembrane regions, aligning with characteristics of G-protein-coupled receptors (GPCRs). The
presence of conserved glycosylation sites and disulfide bonds underscores their roles in protein stability and signal trans-
duction, crucial for maintaining cellular health and promoting longevity.

Analysis of propeptide cleavage sites indicates variability in processing mechanisms among mth proteins, with most pro-
teins predicted to have signal peptide cleavage sites, while some lack these features. This variability suggests diverse

13
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
functional roles and maturation processes. Intrinsic protein disorder analysis reveals significant variability in flexibility,
with some proteins displaying high disorder and others showing more ordered structures. This flexibility is likely impor-
tant for dynamic interactions and signaling, contributing to cellular health and longevity [95].

The analysis of secondary structure, relative solvent-accessible area (RSA), and dihedral angles ϕ and ψ revealed three
group of mth proteins, reflecting their evolutionary relationships and structural similarities.

The study highlights the balance between evolutionary conservation and diversification among mth proteins, underscor-
ing their significance in longevity. The quantitative and comprehensive approach used provides new insights into their
functional versatility and evolutionary dynamics, advancing our understanding of mth proteins in aging and longevity
[96].

Author contributions statement

AG and SSH conceived the problem and theoretical experiments. SSH, DN, VNU, AnG, MS, and PB executed the
results and performed the analysis. SSH, DN, VNU, and KL wrote the initial draft. All authors reviewed and edited the
manuscript. All the authors checked, reviewed, and approved the final version of the manuscript.

Declaration of competing interest

The authors declare no conflict of interest.

References

[1] M. H. Karol, How environmental agents influence the aging process, Biomolecules & Therapeutics 17 (2) (2009)
113–124.
[2] I. Semsei, On the nature of aging, Mechanisms of ageing and development 117 (1-3) (2000) 93–108.
[3] P. D’Aquila, G. Rose, D. Bellizzi, G. Passarino, Epigenetics and aging, Maturitas 74 (2) (2013) 130–136.
[4] P. Davalli, T. Mitic, A. Caporali, A. Lauriola, D. D’Arca, Ros, cell senescence, and novel molecular mechanisms in
aging and age-related diseases, Oxidative medicine and cellular longevity 2016 (1) (2016) 3565127.
[5] J. Aunan, M. Watson, H. Hagland, K. Søreide, Molecular and biological hallmarks of ageing, Journal of British
Surgery 103 (2) (2016) e29–e46.
[6] L. Patanè, R. Strauss, P. Arena, L. Patanè, R. Strauss, P. Arena, Biological investigation of neural circuits in the
insect brain, Nonlinear Circuits and Systems for Neuro-inspired Robot Control (2018) 1–20.
[7] K.-F. Chen, D. C. Crowther, Functional genomics in drosophila models of human disease, Briefings in functional
genomics 11 (5) (2012) 405–415.
[8] A. Ahmed, L. Song, E. P. Xing, Time-varying networks: Recovering temporally rewiring genetic networks during the
life cycle of drosophila melanogaster, arXiv preprint arXiv:0901.0138 (2008).
[9] C. Wan, C. Wan, Background on biology of ageing and bioinformatics, Hierarchical Feature Selection for Knowledge
Discovery: Application of Data Mining to the Biology of Ageing (2019) 25–43.
[10] W. Mair, A. Dillin, Aging and survival: the genetics of life span extension by dietary restriction, Annu. Rev. Biochem.
77 (1) (2008) 727–754.

[11] M. Kaeberlein, K. T. Kirkland, S. Fields, B. K. Kennedy, Genes determining yeast replicative life span in a long-lived
genetic background, Mechanisms of ageing and development 126 (4) (2005) 491–504.
[12] T. B. Kirkwood, R. Holliday, The evolution of ageing and longevity, Proceedings of the Royal Society of London.
Series B. Biological Sciences 205 (1161) (1979) 531–546.
[13] T. Kirkwood, Understanding ageing from an evolutionary perspective, Journal of internal medicine 263 (2) (2008)
117–127.
[14] L. Guarente, C. Kenyon, Genetic pathways that regulate ageing in model organisms, Nature 408 (6809) (2000)
255–262.
[15] J. Vijg, Y. Suh, Genetics of longevity and aging, Annu. Rev. Med. 56 (1) (2005) 193–212.

14
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
[16] C. W. Fox, K. L. Scheibly, W. G. Wallin, L. J. Hitchcock, R. C. Stillwell, B. P. Smith, The genetic architecture of
life span and mortality rates: gender and species differences in inbreeding load of two seed-feeding beetles, Genetics
174 (2) (2006) 763–773.
[17] A. B. Paaby, P. S. Schmidt, Dissecting the genetics of longevity in drosophila melanogaster, Fly 3 (1) (2009) 29–38.

[18] Nisha, K. Raj, Pragati, S. Tandon, S. I. Chanu, S. Sarkar, Aging: Reading, reasoning, and resolving using drosophila
as a model system, Models, Molecules and Mechanisms in Biogerontology: Cellular Processes, Metabolism and
Diseases (2020) 259–302.
[19] S. D. Pletcher, A. A. Khazaeli, J. W. Curtsinger, Why do life spans differ? partitioning mean longevity differences
in terms of age-specific mortality parameters, The Journals of Gerontology Series A: Biological Sciences and Medical
Sciences 55 (8) (2000) B381–B389.
[20] F. A. Lagunas-Rangel, G protein-coupled receptors that influence lifespan of human and animal models, Biogeron-
tology 23 (1) (2022) 1–19.
[21] A. R. Araújo, M. Reis, H. Rocha, B. Aguiar, R. Morales-Hojas, S. Macedo-Ribeiro, N. A. Fonseca, D. Reboiro-Jato,
M. Reboiro-Jato, F. Fdez-Riverola, et al., The drosophila melanogaster methuselah gene: a novel gene with ancient
functions, PloS one 8 (5) (2013) e63747.
[22] A. B. Paaby, P. S. Schmidt, Functional significance of allelic variation at methuselah, an aging gene in drosophila,
PLoS One 3 (4) (2008) e1987.

[23] W. J. William, A. P. West Jr, S. L. Delker, P. J. Bjorkman, S. Benzer, R. W. Roberts, Peptide ligands for methuselah,
a drosophila g protein-coupled receptor associated with extended lifespan, Peptide Modulators of G Protein Signaling
(2005) 118.
[24] S. J. Broughton, M. D. Piper, T. Ikeya, T. M. Bass, J. Jacobson, Y. Driege, P. Martinez, E. Hafen, D. J. Withers,
S. J. Leevers, et al., Longer lifespan, altered metabolism, and stress resistance in drosophila from ablation of cells
making insulin-like ligands, Proceedings of the National Academy of Sciences 102 (8) (2005) 3105–3110.
[25] L. E. Gimenez, P. Ghildyal, K. E. Fischer, H. Hu, W. W. Ja, B. A. Eaton, Y. Wu, S. N. Austad, R. Ranjan, Modulation
of methuselah expression targeted to d rosophila insulin-producing cells extends life and enhances oxidative stress
resistance, Aging cell 12 (1) (2013) 121–129.
[26] Y.-J. Lin, L. Seroude, S. Benzer, Extended life-span and stress resistance in the drosophila mutant methuselah,
Science 282 (5390) (1998) 943–946.
[27] M. D. Adams, S. E. Celniker, R. A. Holt, C. A. Evans, J. D. Gocayne, P. G. Amanatides, S. E. Scherer, P. W.
Li, R. A. Hoskins, R. F. Galle, et al., The genome sequence of drosophila melanogaster, Science 287 (5461) (2000)
2185–2195.
[28] P. S. Schmidt, D. D. Duvernell, W. F. Eanes, Adaptive evolution of a candidate gene for aging in drosophila,
Proceedings of the National Academy of Sciences 97 (20) (2000) 10861–10865.
[29] A. G. Clark, M. B. Eisen, D. R. Smith, C. M. Bergman, B. Oliver, T. A. Markow, T. C. Kaufman, M. Kellis,
W. Gelbart, V. N. Iyer, et al., Evolution of genes and genomes on the drosophila phylogeny, Nature 450 (7167)
(2007) 203–18.

[30] D. D. Duvernell, P. S. Schmidt, W. F. Eanes, Clines and adaptive evolution in the methuselah gene region in
drosophila melanogaster, Molecular ecology 12 (5) (2003) 1277–1285.
[31] F. Sievers, D. G. Higgins, Clustal omega, Current protocols in bioinformatics 48 (1) (2014) 3–13.
[32] F. Sievers, A. Wilm, D. Dineen, T. J. Gibson, K. Karplus, W. Li, R. Lopez, H. McWilliam, M. Remmert, J. Söding,
et al., Fast, scalable generation of high-quality protein multiple sequence alignments using clustal omega, Molecular
systems biology 7 (1) (2011) 539.
[33] S. Kumar, G. Stecher, M. Li, C. Knyaz, K. Tamura, Mega x: molecular evolutionary genetics analysis across com-
puting platforms, Molecular biology and evolution 35 (6) (2018) 1547–1549.
[34] S. S. Hassan, D. Attrish, S. Ghosh, P. P. Choudhury, V. N. Uversky, A. A. Aljabali, K. Lundstrom, B. D. Uhal,
N. Rezaei, M. Seyran, et al., Notable sequence homology of the orf10 protein introspects the architecture of sars-cov-2,
International Journal of Biological Macromolecules 181 (2021) 801–809.
[35] S. S. Hassan, P. P. Choudhury, P. Basu, S. S. Jana, Molecular conservation and differential mutation on orf3a gene
in indian sars-cov2 genomes, Genomics 112 (5) (2020) 3226–3237.

15
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
[36] D. Nawn, S. S. Hassan, M. Sil, A. Ghosh, A. Goswami, P. Basu, G. W. Dayhoff II, K. Lundstrom, V. N. Uversky,
The distal-proximal relationships among the human moonlighting proteins: Evolutionary hotspots and darwinian
checkpoints, International Journal of Biological Macromolecules 259 (2024) 128998.
[37] R. P. Trees, The neighbor-joining method: a new method for, Mol Biol Evol 4 (4) (1987) 406–425.

[38] K. Tamura, G. Stecher, S. Kumar, Mega11: molecular evolutionary genetics analysis version 11, Molecular biology
and evolution 38 (7) (2021) 3022–3027.
[39] J. D. Chavez, C. R. Weisbrod, C. Zheng, J. K. Eng, J. E. Bruce, Protein interactions, post-translational modifications
and topologies in human cells, Molecular & Cellular Proteomics 12 (5) (2013) 1451–1467.

[40] P. Duckert, S. Brunak, N. Blom, Prediction of proprotein convertase cleavage sites, Protein Engineering Design and
Selection 17 (1) (2004) 107–112.
[41] W.-L. Hsu, C. Oldfield, J. Meng, F. Huang, B. Xue, V. N. Uversky, P. Romero, A. K. Dunker, Intrinsic protein
disorder and protein-protein interactions, in: Biocomputing 2012, World Scientific, 2012, pp. 116–127.

[42] J. Habchi, P. Tompa, S. Longhi, V. N. Uversky, Introducing protein intrinsic disorder, Chemical reviews 114 (13)
(2014) 6561–6588.
[43] V. N. Uversky, Protein intrinsic disorder and structure-function continuum, Progress in molecular biology and trans-
lational science 166 (2019) 1–17.
[44] A. Bhattarai, I. A. Emerson, Dynamic conformational flexibility and molecular interactions of intrinsically disordered
proteins, Journal of biosciences 45 (1) (2020) 29.
[45] V. N. Uversky, A decade and a half of protein intrinsic disorder: biology still waits for physics, Protein Science 22 (6)
(2013) 693–724.
[46] B. Xue, R. L. Dunbrack, R. W. Williams, A. K. Dunker, V. N. Uversky, Pondr-fit: a meta-predictor of intrinsically
disordered amino acids, Biochimica et Biophysica Acta (BBA)-Proteins and Proteomics 1804 (4) (2010) 996–1010.
[47] M. H. Høie, E. N. Kiehl, B. Petersen, M. Nielsen, O. Winther, H. Nielsen, J. Hallgren, P. Marcatili, Netsurfp-3.0:
accurate and fast prediction of protein structural features by protein language models and deep learning, Nucleic
acids research 50 (W1) (2022) W510–W515.
[48] E. V. Koonin, Orthologs, paralogs, and evolutionary genomics, Annu. Rev. Genet. 39 (1) (2005) 309–338.
[49] J. Zhang, Evolution by gene duplication: an update, Trends in ecology & evolution 18 (6) (2003) 292–298.
[50] I. S. Moreira, P. A. Fernandes, M. J. Ramos, Hot spots—a review of the protein–protein interface determinant
amino-acid residues, Proteins: Structure, Function, and Bioinformatics 68 (4) (2007) 803–812.
[51] M. J. Page, E. Di Cera, Serine peptidases: classification, structure and function, Cellular and Molecular Life Sciences
65 (2008) 1220–1236.
[52] M. Shatsky, R. Nussinov, H. J. Wolfson, A method for simultaneous alignment of multiple protein structures, Proteins:
Structure, Function, and Bioinformatics 56 (1) (2004) 143–156.
[53] H. Cid, M. Bunster, M. Canales, F. Gazitúa, Hydrophobicity and structural classes in proteins, Protein Engineering,
Design and Selection 5 (5) (1992) 373–375.
[54] S. Khemaissa, S. Sagan, A. Walrant, Tryptophan, an amino-acid endowed with unique properties and its many roles
in membrane proteins, Crystals 11 (9) (2021) 1032.
[55] E. S. Epel, G. J. Lithgow, Stress biology and aging mechanisms: toward understanding the deep connection between
adaptation to stress and longevity, Journals of Gerontology Series A: Biomedical Sciences and Medical Sciences
69 (Suppl 1) (2014) S10–S16.
[56] C. Kenyon, The first long-lived mutants: discovery of the insulin/igf-1 pathway for ageing, Philosophical Transactions
of the Royal Society B: Biological Sciences 366 (1561) (2011) 9–16.
[57] C. Li, Y. Zhang, X. Yun, Y. Wang, M. Sang, X. Liu, X. Hu, B. Li, M ethuselah-like genes affect development, stress
resistance, lifespan and reproduction in t ribolium castaneum, Insect molecular biology 23 (5) (2014) 587–597.

[58] C. A. Elena-Real, P. Mier, N. Sibille, M. A. Andrade-Navarro, P. Bernadó, Structure–function relationships in protein


homorepeats, Current Opinion in Structural Biology 83 (2023) 102726.
[59] C. López-Otı́n, M. A. Blasco, L. Partridge, M. Serrano, G. Kroemer, The hallmarks of aging, Cell 153 (6) (2013)
1194–1217.

16
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
[60] L. Fontana, L. Partridge, V. D. Longo, Extending healthy life span—from yeast to humans, science 328 (5976) (2010)
321–326.
[61] V. Tiwari, S. D. Karpe, R. Sowdhamini, Topology prediction of insect olfactory receptors, Current Opinion in
Structural Biology 55 (2019) 194–203.
[62] D. Calebiro, Z. Koszegi, The subcellular dynamics of gpcr signaling, Molecular and cellular endocrinology 483 (2019)
24–30.
[63] P. M. Couto, J. J. Caramelo, Glycoprotein folding, in: Molecular Nutrition: Carbohydrates, Elsevier, 2019, pp.
59–71.
[64] N. G. Jayaprakash, A. Surolia, Role of glycosylation in nucleating protein folding and stability, Biochemical Journal
474 (14) (2017) 2333–2347.
[65] G. Bulaj, Formation of disulfide bonds in proteins and peptides, Biotechnology advances 23 (1) (2005) 87–92.
[66] R. Fredriksson, M. C. Lagerström, L.-G. Lundin, H. B. Schiöth, The g-protein-coupled receptors in the human genome
form five main families. phylogenetic analysis, paralogon groups, and fingerprints, Molecular pharmacology 63 (6)
(2003) 1256–1272.
[67] M. Berger, M. Kaup, V. Blanchard, Protein glycosylation and its impact on biotechnology, Genomics and Systems
Biology of Mammalian Cell Culture (2012) 165–185.
[68] E. Nisoli, E. Clementi, S. Moncada, M. O. Carruba, Mitochondrial biogenesis as a cellular signaling framework,
Biochemical pharmacology 67 (1) (2004) 1–15.
[69] C. S. Odoemelam, B. Percival, H. Wallis, M.-W. Chang, Z. Ahmad, D. Scholey, E. Burton, I. H. Williams, C. L.
Kamerlin, P. B. Wilson, G-protein coupled receptors: structure and function in drug discovery, RSC advances 10 (60)
(2020) 36337–36348.
[70] M. Von Zastrow, Role of endocytosis in signalling and regulation of g-protein-coupled receptors, Biochemical Society
Transactions 29 (4) (2001) 500–504.
[71] M. V. Trivedi, J. S. Laurence, T. J. Siahaan, The role of thiols and disulfides on protein stability, Current Protein
and Peptide Science 10 (6) (2009) 614–625.
[72] S.-O. Yoon, C.-H. Yun, A.-S. Chung, Dose effect of oxidative stress on signal transduction in aging, Mechanisms of
ageing and development 123 (12) (2002) 1597–1604.
[73] K. H. Choo, S. Ranganathan, Flanking signal and mature peptide residues influence signal peptide cleavage, BMC
bioinformatics 9 (2008) 1–11.
[74] G. Palade, Intracellular aspects of the process of protein synthesis, Science 189 (4200) (1975) 347–358.
[75] H. Owji, N. Nezafat, M. Negahdaripour, A. Hajiebrahimi, Y. Ghasemi, A comprehensive review of signal peptides:
Structure, roles, and applications, European journal of cell biology 97 (6) (2018) 422–441.
[76] S. Özöğür-Akyüz, J. Shawe-Taylor, G.-W. Weber, Z. Ögel, Pattern analysis for the prediction of fungal pro-peptide
cleavage sites, Discrete applied mathematics 157 (10) (2009) 2388–2394.
[77] D. J. Klionsky, S. D. Emr, Membrane protein sorting: biosynthesis, transport and processing of yeast vacuolar
alkaline phosphatase., The EMBO journal 8 (8) (1989) 2241–2250.
[78] T. Klein, U. Eckhard, A. Dufour, N. Solis, C. M. Overall, Proteolytic cleavage mechanisms, function, and “omic”
approaches for a near-ubiquitous posttranslational modification, Chemical reviews 118 (3) (2018) 1137–1168.
[79] P. Walter, A. E. Johnson, Signal sequence recognition and protein targeting to the endoplasmic reticulum membrane,
Annual review of cell biology 10 (1) (1994) 87–119.
[80] A. Capell, H. Steiner, M. Willem, H. Kaiser, C. Meyer, J. Walter, S. Lammich, G. Multhaup, C. Haass, Maturation
and pro-peptide cleavage of β-secretase, Journal of Biological Chemistry 275 (40) (2000) 30849–30854.
[81] I. Bludau, R. Aebersold, Proteomic and interactomic insights into the molecular basis of cell functional diversity,
Nature Reviews Molecular Cell Biology 21 (6) (2020) 327–340.
[82] J. Pei, L. N. Kinch, Q. Cong, Computational analysis of propeptide-containing proteins and prediction of their
post-cleavage conformation changes, Proteins: Structure, Function, and Bioinformatics (2024).
[83] B. Cauwe, G. Opdenakker, Intracellular substrate cleavage: a novel dimension in the biochemistry, biology and
pathology of matrix metalloproteinases, Critical reviews in biochemistry and molecular biology 45 (5) (2010) 351–
423.

17
bioRxiv preprint doi: [Link] this version posted November 6, 2024. The copyright holder for this
preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in
perpetuity. It is made available under aCC-BY-ND 4.0 International license.
[84] R. Trivedi, H. A. Nagarajaram, Intrinsically disordered proteins: an overview, International journal of molecular
sciences 23 (22) (2022) 14050.
[85] S. Naskar, G. R. Gowane, A. Chopra, C. Paswan, L. L. L. Prince, Genetic adaptability of livestock to environmental
stresses, Environmental stress and amelioration in livestock production (2012) 317–378.

[86] B. Henderson, M. A. Fares, A. C. Martin, Protein moonlighting in biology and medicine, John Wiley & Sons, 2016.
[87] L. Wang, J. Karpac, H. Jasper, Promoting longevity by maintaining metabolic and proliferative homeostasis, Journal
of Experimental Biology 217 (1) (2014) 109–118.
[88] R. Adamczak, A. Porollo, J. Meller, Combining prediction of secondary structure and solvent accessibility in proteins,
Proteins: Structure, Function, and Bioinformatics 59 (3) (2005) 467–475.
[89] B. Rost, C. Sander, Conservation and prediction of solvent accessibility in protein families, Proteins: Structure,
Function, and Bioinformatics 20 (3) (1994) 216–226.
[90] L. S. Swapna, S. Mahajan, A. G. de Brevern, N. Srinivasan, Comparison of tertiary structures of proteins in protein-
protein complexes with unbound forms suggests prevalence of allostery in signalling proteins, BMC Structural Biology
12 (2012) 1–21.
[91] J. Felsenstein, Phylogenies and the comparative method, The American Naturalist 125 (1) (1985) 1–15.
[92] D. L. Swofford, J. Sullivan, et al., Phylogeny inference based on parsimony and other methods using paup*, The
phylogenetic handbook: a practical approach to DNA and protein phylogeny 7 (2003) 160–206.

[93] B. K. Kennedy, D. W. Lamming, The mechanistic target of rapamycin: the grand conductor of metabolism and
aging, Cell metabolism 23 (6) (2016) 990–1003.
[94] H. M. Berman, J. Westbrook, Z. Feng, G. Gilliland, T. N. Bhat, H. Weissig, I. N. Shindyalov, P. E. Bourne, The
protein data bank, Nucleic acids research 28 (1) (2000) 235–242.
[95] A. K. Dunker, C. J. Oldfield, J. Meng, P. Romero, J. Y. Yang, J. W. Chen, V. Vacic, Z. Obradovic, V. N. Uversky,
The unfoldomics decade: an update on intrinsically disordered proteins, BMC genomics 9 (2008) 1–26.
[96] L. Fontana, L. Partridge, Promoting health and longevity through diet: from model organisms to humans, Cell
161 (1) (2015) 106–118.

18

You might also like