0% ont trouvé ce document utile (0 vote)
3 vues206 pages

Key Principles of Graph Machine Learning - Representation

Cette thèse examine les défis des réseaux de neurones de graphes (GNNs) en termes de généralisation, de robustesse et d'apprentissage de représentations. Elle propose trois contributions principales : de nouvelles techniques d'apprentissage basées sur des opérateurs de décalage de graphe, des méthodes d'augmentation de données pour améliorer la généralisation, et des GNNs plus robustes grâce à des techniques d'orthonormalisation. Ce travail offre une compréhension approfondie des limitations et du potentiel des GNNs, tout en fournissant des solutions pratiques pour améliorer leur performance.

Transféré par

marisageek
Copyright
© All Rights Reserved
Nous prenons très au sérieux les droits relatifs au contenu. Si vous pensez qu’il s’agit de votre contenu, signalez une atteinte au droit d’auteur ici.
Formats disponibles
Téléchargez aux formats PDF, TXT ou lisez en ligne sur Scribd
0% ont trouvé ce document utile (0 vote)
3 vues206 pages

Key Principles of Graph Machine Learning - Representation

Cette thèse examine les défis des réseaux de neurones de graphes (GNNs) en termes de généralisation, de robustesse et d'apprentissage de représentations. Elle propose trois contributions principales : de nouvelles techniques d'apprentissage basées sur des opérateurs de décalage de graphe, des méthodes d'augmentation de données pour améliorer la généralisation, et des GNNs plus robustes grâce à des techniques d'orthonormalisation. Ce travail offre une compréhension approfondie des limitations et du potentiel des GNNs, tout en fournissant des solutions pratiques pour améliorer leur performance.

Transféré par

marisageek
Copyright
© All Rights Reserved
Nous prenons très au sérieux les droits relatifs au contenu. Si vous pensez qu’il s’agit de votre contenu, signalez une atteinte au droit d’auteur ici.
Formats disponibles
Téléchargez aux formats PDF, TXT ou lisez en ligne sur Scribd

NNT : 2025IPPAX053

Key Principles of Graph Machine


Learning: Representation, Robustness,
and Generalization
Thèse de doctorat de l’Institut Polytechnique de Paris
arXiv:2602.01139v1 [[Link]] 1 Feb 2026

préparée à l’École Polytechnique

École doctorale n◦ 626 : l’École Doctorale de l’Institut Polytechnique de Paris


(ED IP Paris)
Spécialité de doctorat: Mathématiques et Informatique

Thèse présentée et soutenue à Palaiseau, le 9 juillet 2025, par YASSINE

A BBAHADDOU

Composition du Jury :

Céline Hudelot
Professor, CentraleSupélec Presidente, Examinatrice

Charlotte Laclau
Professor, Télécom Paris Examinatrice

Davide Bacciu
Professor, Università di Pisa Examinateur

Thomas Gärtner
Professor, Technische Rapporteur
Universität Wien
Marc Lelarge
Researcher, INRIA Rapporteur

Michalis Vazirgiannis
Full Professor, École Directeur de thèse
Polytechnique (LIX)
Fragkiskos D. Malliaros
Associate Professor, Co-encadrant de thèse
CentraleSupelec (CVN)
Johannes F. Lutzeyer
626 : Assistant Professor, École Co-encadrant de thèse
Polytechnique (LIX)
ABSTRACT

Graph Neural Networks (GNNs) have emerged as powerful tools for learning representations from
structured data. Despite their growing popularity and success across various applications, GNNs encounter
several challenges that limit their performance. in their generalization, robustness to adversarial perturba-
tions, and the effectiveness of their representation learning capabilities. In this dissertation, I investigate
these core aspects through three main contributions: (1) developing new representation learning tech-
niques based on Graph Shift Operators (GSOs, aiming for enhanced performance across various contexts
and applications, (2) introducing generalization-enhancing methods through graph data augmentation,
and (3) developing more robust GNNs by leveraging orthonormalization techniques and noise-based
defenses against adversarial attacks. By addressing these challenges, my work provides a more principled
understanding of the limitations and potential of GNNs.

ii
RÉSUMÉ

Les réseaux de neurones de graphes (GNNs) sont devenus des outils puissants pour l’apprentissage
de représentations à partir de données structurées. Malgré leur popularité croissante et leur succès dans
diverses applications, les GNNs rencontrent plusieurs défis qui limitent leurs performances, notamment en
termes de généralisation, de robustesse aux perturbations adversariales. Dans cette thèse, j’examine ces
aspects fondamentaux à travers trois principales contributions.
La première contribution se concentre sur l’amélioration de l’apprentissage des représentations. J’intro-
duis une nouvelle famille d’opérateurs de décalage de graphe (GSOs) appelés Opérateurs de Décalage
de Graphe basés sur la Centralité (CGSOs). Contrairement aux GSOs traditionnels qui se basent princi-
palement sur des informations locales comme le degré des nœuds, les CGSOs intègrent des métriques
de centralité globale, telles que le PageRank, le k-core, pour normaliser la matrice d’adjacence. Cette
approche enrichit la représentation du graphe avec une information structurelle globale tout en préservant
sa sparsité, un atout majeur pour l’efficacité calculatoire sur de grands réseaux. Une analyse spectrale
approfondie de ces opérateurs révèle leurs propriétés fondamentales et leur impact sur la diffusion de
l’information. De plus, je propose une seconde méthode, ADMP-GNN, qui s’attaque à la limitation des
GNNs standards utilisant le même nombre d’operations pour tous les nœuds. En démontrant empirique-
ment que la profondeur optimale varie selon les caractéristiques locales des nœuds, cet approche ajuste
dynamiquement le nombre de couches de propagation pour chaque nœud, améliorant ainsi la précision
prédictive.
La deuxième contribution vise à améliorer la généralisation des GNNs, en particulier pour la classification
de graphes face à des données limitées ou hors-distribution (OOD). Je développe d’abord un cadre
théorique basé sur la complexité de Rademacher pour analyser l’impact de l’augmentation de données
sur l’erreur de généralisation. Ce cadre montre que la maîtrise de la distance entre les représentations
originales et augmentées est cruciale. Sur cette base, je propose GRATIN, une nouvelle technique
d’augmentation qui opère dans l’espace latent des représentations de graphes. En utilisant des Modèles
de Mélange Gaussien (GMMs) pour modéliser la distribution des graphes de chaque classe, GRATIN
génère de nouveaux échantillons synthétiques qui enrichissent l’ensemble d’entraînement. Cette approche
se distingue par son efficacité calculatoire et sa capacité à améliorer significativement la performance de
généralisation.
La troisième contribution porte sur le développement de GNNs plus robustes. D’une part, j’établis un
lien théorique entre la robustesse attendue d’un GNN et l’orthonormalité de ses matrices de poids, ce qui
conduit à la proposition de GCORN (Graph Convolutional Orthonormal Robust Networks). Ce modèle, une
variante robuste du GCN, intègre une contrainte d’orthonormalisation durant l’entraînement pour améliorer
sa résistance aux attaques sur les attributs des nœuds. D’autre part, je propose RobustCRF, une méthode
de défense post-hoc et agnostique du modèle, qui opère au stade de l’inférence. Cette approche innovante
peut ainsi renforcer n’importe quel GNN pré-entraîné sans modification architecturale. En s’appuyant sur
un Champ Aléatoire Conditionnel (CRF), cette technique affine les prédictions du modèle en exploitant
l’hypothèse que des graphes structurellement similaires dans le voisinage d’une instance donnée devraient
produire des prédictions similaires, corrigeant ainsi les erreurs induites par des perturbations adverses
sans nécessiter un réentraînement.
En relevant ces défis, mon travail apporte une compréhension plus rigoureuse des limites et du potentiel
des GNNs, tout en proposant des solutions pratiques et théoriquement fondées pour améliorer leur
performance, leur généralisation et leur robustesse.

iii
Au-delà des contributions méthodologiques, la thèse insiste sur la reproductibilité, des analyses d’ablation
et une étude empirique sur des jeux de données hétérogènes (nœuds/graphes) pour éclairer les compromis
expressivité–coût. Elle ouvre des pistes en bio-informatique, recommandation et sécurité, et discute des
questions éthiques liées à la manipulation de graphes.

iv
P U B L I C AT I O N S

Several contributions presented in this dissertation relate to the following peer-reviewed articles:
— Abbahaddou, Y., Ennadir, S., Lutzeyer, J. F., Vazirgiannis, M., & Boström, H. (2024). Bounding the
Expected Robustness of Graph Neural Networks Subject to Node Feature Attacks. In The Twelfth
International Conference on Learning Representations (ICLR).
— Abbahaddou, Y., Malliaros, F. D., Lutzeyer, J. F., Aboussalah, A. M., & Vazirgiannis, M. (2025).
Graph Neural Network Generalization With Gaussian Mixture Model Based Augmentation. In The
Forty-Second International Conference on Machine Learning (ICML).
— Abbahaddou, Y., Ennadir, S., Lutzeyer, J. F., Vazirgiannis, M., & Malliaros, F. D. (2024). Rethinking
Robustness in Graph Neural Networks: A Post-Hoc Approach With Conditional Random Fields. arXiv
preprint arXiv:2411.05399.
— Abbahaddou, Y., Malliaros, F. D., Vazirgiannis, M., & Lutzeyer, J. F. (2024). Centrality Graph Shift
Operators for Graph Neural Networks. arXiv preprint arXiv:2411.04655.
— Abbahaddou, Y., Malliaros, F. D., Lutzeyer, J. F., Aboussalah, A. M., & Vazirgiannis, M. (2024).
Gaussian Mixture Models Based Augmentation Enhances GNN Generalization. arXiv preprint
arXiv:2411.08638.
— Abbahaddou, Y., Ennadir, S., Lutzeyer, J. F., Malliaros, F. D., & Vazirgiannis, M. (2024). Post-Hoc
Robustness Enhancement in Graph Neural Networks with Conditional Random Fields. arXiv preprint
arXiv:2411.05399.

The following publication is related, but is not extensively discussed in this thesis:
— Ennadir, S., Abbahaddou, Y., Lutzeyer, J. F., Vazirgiannis, M., & Boström, H. (2024). A Simple and
Yet Fairly Effective Defense for Graph Neural Networks. Proceedings of the AAAI Conference on
Artificial Intelligence, 38(19), 21063–21071.
— Abbahaddou, Y., Lutzeyer, J., & Vazirgiannis, M. (2023). Graph Neural Networks on Discriminative
Graphs of Words. In NeurIPS 2023 Workshop: New Frontiers in Graph Learning.
— Ennadir, S., Smirnov, O., Abbahaddou, Y., Cao, L., & Lutzeyer, J. F. (2025). Enhancing Graph
Classification Robustness with Singular Pooling. In The Thirty-ninth Annual Conference on Neural
Information Processing Systems (NeurIPS).
— Abbahaddou, Y., & Aboussalah, A. M. (2025). A Geometry-Aware Metric for Mode Collapse in
Multivariate Time Series Generative Models. In The Thirty-ninth Annual Conference on Neural
Information Processing Systems (NeurIPS).

v
vi
LIST OF ACRONYMS AND SYMBOLS

Common Acronyms in Alphabetic Order


— AMI: Adjusted Mutual Information;
— ARI: Adjusted Rand Index;
— BA: Barabási–Albert Model;
— GAT: Graph Attention Networks;
— GATv2: Graph Attention Networks v2;
— GCN: Graph Convolutional Network;
— GIN: Graph Isomorphism Network;
— GNN: Graph Neural Network;
— GSO: Graph Shift Operator;
— GSP: Graph Signal Processing;
— PNA: Principal Neighbourhood Aggregation;
— SBM: Stochastic Block Model;
Common Linear Algebra Symbols and Operations For some dimension p ∈ N∗ , some vectors a ∈ R p
and b ∈ R p , and some square matrices N ∈ R p× p and M ∈ R p× p :
— 1 p or 1 : The p−dimensional vector containing p entries all equal to 1;
— ai ∈ R : The i −th element of a, for some i ∈ {1, . . . , p};
— a⊤ : The transposition of the vector a;
— mi,j ∈ R, Mi,j ∈ R, or M(i, j) ∈ R the (i, j)−th element of M, for some i, j ∈ {1, . . . , p};
Common Graph Representation Learning Symbols
— G = (V , E ) : A graph composed of a node set V and an edge set E ;
— i ∈ V : A node i from G ;
— (i, j) ∈ E : An edge connecting node i to node j in G ;
— n ∈ N∗ : The number of nodes in G , i.e., n = |V |;
— m ∈ N: The number of edges in G , i.e., m = |E |;
— N (i ) ⊆ V: The neighborhood of a node i ∈ V ;
— I ∈ Rn×n : The Identity matrix;
— A ∈ [0, 1]n×n : The Adjacency matrix of G ;
— D ∈ Rn×n : The degree matrix, i.e., a diagonal matrix defined as Dii = ∑in=1 aij ;
e ∈ Rn×n : The normalized adjacency matrix of G , i.e., A
— A e = D−1/2 AD−1/2 ;
— L ∈ Rn×n : The Unnormalised Laplacian matrix of G ;
— Q ∈ Rn×n : The Signless Laplacian matrix of G ;
— Lrw ∈ Rn×n : The Random-walk Normalised Laplacian of G ;

vii
— Lsym ∈ Rn×n : The Symmetric Normalised Laplacian matrix of G ;
— Â ∈ Rn×n : The Normalised Adjacency matrix of G ;
— λ1 (Lrw ) : The Equidistribution radius of G ;
— ϱr ( H ) : The Normalized Spectral Gap of G ;
— XG ∈ Rn×K , or simply X ∈ Rn×K : The node feature matrix of the graph G , obtained by stacking up all
xi vectors;
— Edir (·) : The Dirichlet Energy function;
— deg(i ) : The degree of a node i;
— deg+ (i ) : The out-degree of a node i;
— deg− (i ) : The in-degree of a node i;
— δ(G) : The minimum degree of the graph G ;
— ∆(G) : The maximum degree of the graph G ;
— hC (G) : The Cheeger constant of the graph G ;
— γh : The Homophily ratio of the graph G ;
— G : A set of multiple graphs;
— Gtrain : The set of training graphs;
— Gval : The set of validation graphs;
— Gtest : The set of test graphs;
— L : The number of message-passing layers in the GNN;
— Vtrain : The set of training nodes;
— Vval : The set of validation nodes;
— Vtest : The set of test nodes;
— Ytrain : The labels of the training nodes Vtrain ⊂ V ;
— Ytest : The labels of the test nodes Vtest ⊂ V ;
— Yval : The labels of the validation nodes Vval ⊂ V ;
— Y : The labels space;
— f : G → Y : A graph function;
— READOUT : The Readout function in the GNN;
— COMBINE : The Combine function in the GNN;
— AGGREGATE : The Aggregate function in the GNN;
— MLP(ℓ) : A multi-layer perceptron used at the ℓ-th layer of the GNN;

viii
CONTENTS

I Introduction
1 Introduction 3
1.1 The Challenges in GNNs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.2 Overview of Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2 Preliminaries 7
2.1 Notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
2.2 Graph Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
2.2.1 Heterogeneity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2.2.2 Homophily . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2.2.3 Sparsity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.2.4 Connectivity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.3 Graph Exploraton: Walks and Random Walks . . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.3.1 Walks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.3.2 Random Walks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.4 Graph Shift Operators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.4.1 Definition of Graph Shift Operators . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.4.2 Spectral Clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
2.4.3 Spectral Properties of GSOs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.5 Node Centrality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
2.5.1 k-core . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
2.5.2 PageRank . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.5.3 Walk Count . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
2.6 Graph Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
2.6.1 Key GNN Architectures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.6.2 Challenges in GNNs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
2.6.3 Generalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
2.6.4 Expressivity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.7 Graph Learning Tasks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.7.1 Semi-Supervised Node classification . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.7.2 Graph Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.8 Graph Distance Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.9 Synthetic Graphs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
2.9.1 Stochastic Block Model (SBM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
2.9.2 Barabási–Albert (BA) Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.10 Benchmarks and Datasets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.10.1 Datasets for the Node Classification Task . . . . . . . . . . . . . . . . . . . . . . . . 23
2.10.2 Datasets for the Graph Classification Task . . . . . . . . . . . . . . . . . . . . . . . . 24
2.11 Evaluation Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
2.11.1 Node Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
2.11.2 Graph Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
2.12 Software and Tools . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
2.12.1 Programming Languages and Frameworks . . . . . . . . . . . . . . . . . . . . . . . 25

ix
x CONTENTS

2.12.2 Graph Processing Libraries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25

IIRepresentation Learning in Graph Neural Networks


3 Rethinking GSOs in GNNs for Graph Representation Learning 29
3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
3.3 CGSO: Centrality Graph Shift Operators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
3.3.1 Mathematical Formulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
3.3.2 CGNN: Centrality Graph Neural Network . . . . . . . . . . . . . . . . . . . . . . . . . 33
3.4 A Spectral Clustering Perspective of CGSOs . . . . . . . . . . . . . . . . . . . . . . . . . . 35
3.4.1 Spectral Clustering on Stochastic Block Barabási–Albert Models . . . . . . . . . . . 35
3.4.2 Centrality Recovery in Spectral Clustering . . . . . . . . . . . . . . . . . . . . . . . . 36
3.5 Experimental Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
3.5.1 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.6 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
4 Adaptive Depth Message Passing GNN 41
4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
4.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4.3 Adaptive Depth Message Passing GNN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
4.3.1 Depth Analysis on Synthetic Graphs . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
4.3.2 Adaptive Message Passing Layer Integration . . . . . . . . . . . . . . . . . . . . . . 44
4.3.3 Training Scheme of ADMP-GNN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
4.3.4 Empirical Insights into Node Specific Depth . . . . . . . . . . . . . . . . . . . . . . . 46
4.3.5 Generalization to Test Nodes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
4.4 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
4.4.1 Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
4.4.2 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
4.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51

Robustness of Graph Neural Networks


III
5 Adversarial Robustness of GNNs: Theory, Bounds, and Defense 55
5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
5.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
5.3 Expected Adversarial Robustness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
5.4 The Expected Robustness of Message Passing Neural Networks . . . . . . . . . . . . . . . 58
5.4.1 On the Expected Robustness of Graph Convolutional Networks . . . . . . . . . . . . 58
5.4.2 On the Generalization to Other Graph Neural Networks . . . . . . . . . . . . . . . . 59
5.4.3 Enhancing the Robustness of Graph Convolutional Networks . . . . . . . . . . . . . 60
5.5 Estimation of Our Robustness Measure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
5.6 Empirical Investigation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
5.6.1 Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
5.6.2 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
5.7 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65
6 A Post-hoc Approach With Conditional Random Fields 67
6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
6.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
6.3 Proposed Method: RobustCRF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
CONTENTS xi

6.3.1 Adversarial Attacks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70


6.3.2 Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
6.3.3 Modeling the Robustness Constraint with a CRF . . . . . . . . . . . . . . . . . . . . 71
6.3.4 Mean Field Approximation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
6.3.5 Reducing the Size of the CRF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
6.4 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
6.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77

IVGeneralization of Graph Neural Networks


7 Improving Generalization in GNNs through Data Augmentation 81
7.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81
7.2 Background and Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82
7.3 GRATIN: Gaussian Mixture Model for Graph Data Augmentation . . . . . . . . . . . . . . . 83
7.3.1 Formalism of Graph Data Augmentation . . . . . . . . . . . . . . . . . . . . . . . . . 83
7.3.2 Proposed Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
7.3.3 Time Complexity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
7.3.4 Analyzing the Generalization Ability of the Augmented Graphs via Influence Functions 88
7.3.5 Fisher-Guided GMM Augmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
7.4 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
7.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91

V Conclusion
8 Conclusion 95
8.1 Summary of contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
8.1.1 Representation learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
8.1.2 Robustness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
8.1.3 Generalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
8.2 Broader Impact . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
8.3 Future directions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97

Bibliography 98

VIAppendices
A Appendix: Rethinking GSOs in GNNs for Graph Representation Learning 119
A .1 Datasets and Implementation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
A .1.1 Statistics of the Node Classification Datasets . . . . . . . . . . . . . . . . . . . . . . 120
A .1.2 Implementation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
A .1.3 Weights Initialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
A .1.4 Hyperparameter Configurations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121
A .2 Additional Results for the Node Classification Task . . . . . . . . . . . . . . . . . . . . . . . 122
A .3 Simple Graph Convolutional Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
A .4 Combining Local and Global Centralities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
A .5 The Graph Structure of the Stochastic Block Barabási–Albert Models . . . . . . . . . . . . . 122
A .6 k-core distribution in Stochastic Block Barabási–Albert Models . . . . . . . . . . . . . . . . 123
A .7 Spectral Clustering Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123
A .8 Additional Results for the Spectral Clustering Task . . . . . . . . . . . . . . . . . . . . . . . 123
A .9 CGNN with Heterophily . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
xii CONTENTS

A .10 Proofs of Propositions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124


A .10.1 Proof of Proposition 3.3.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
A .10.2 Proof of Proposition 3.3.2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
A .10.3 Proof of Proposition 3.3.4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
A .11 Average Degree of a Barabasi–Albert Model . . . . . . . . . . . . . . . . . . . . . . . . . . 134
A .12 Learned Parameters of Different Centrality Based GSOs . . . . . . . . . . . . . . . . . . . . 134
A .12.1 Degree Centrality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135
A .12.2 k-Core Centrality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135
A .12.3 PageRank Centrality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135
A .12.4 Count of Walks Centrality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136
B Appendix: Adaptive Depth Message Passing GNN 139
B .1 Dataset Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140
B .2 Node-Specific Depth Analysis in Graph Neural Networks . . . . . . . . . . . . . . . . . . . . 140
B .3 Time Complexity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140
B .4 Hyperparameter Configurations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
B .5 Supplementary Results of ADMP-GCN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
B .6 Supplementary Results of ADMP-GIN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
C Appendix: Adversarial Robustness of GNNs 147
C.1 Proof Of Lemma 5.3.2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148
C.2 Proof Of Proposition c.2.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148
C.3 Proof Of Theorem 5.4.1 and 5.4.2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150
C.4 Proof Of Theorem 5.4.3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152
C.5 Generalizing to Any Graph Neural Network . . . . . . . . . . . . . . . . . . . . . . . . . . . 154
C.6 Vulnerability Upper-bound When Dealing With Both Structural and Node Features Attacks 154
C.6.1 Experimental Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155
C.7 Time and Complexity Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155
C.7.1 On the Effect Of Order/Iterations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155
C.7.2 On Training Time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156
C.8 More Details About The Estimation Of Our Robustness Measure . . . . . . . . . . . . . . . 156
C.8.1 Case Where p < ∞ . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157
C.8.2 Case Where p = ∞ . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 158
C.9 Proof of Lemma 5.5.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159
C.10 Additional Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
C.10.1 Node Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
C.10.2 Graph Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
C.11 Experimental Results on GIN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161
C.12 Datasets and Implementation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161
C.12.1 Node Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161
C.12.2 Graph Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162
C.12.3 Implementation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163
C.12.4 Implementation Details of our Empirical Robustness Estimation . . . . . . . . . . . . 163
D Appendix: A Post-hoc Approach With Conditional Random Fields 165
D.1 Proof of Lemma 6.3.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166
D.2 Asymptotic Behavior of CRF Neighborhood Size . . . . . . . . . . . . . . . . . . . . . . . . 167
D.2.1 Proof of Lemma 6.3.3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
D.2.2 Empirical Investigation of the Lower Bound . . . . . . . . . . . . . . . . . . . . . . . 169
D.3 The Set of Hyperparameters Used to Construct the CRF . . . . . . . . . . . . . . . . . . . . 169
CONTENTS xiii

D.4 Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170


D.5 Time and Complexity of RobustCRF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170
D.6 Resuls on OGBN-Arxiv . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170
E Appendix: Improving Generalization in GNNs through Data Augmentation 173
E .1 Proof of Theorem 7.3.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174
E .2 Proof of Proposition 7.3.2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 176
E .3 Proof of Theorem 7.3.4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 178
E .4 Mathematical Expressions of GCN and GIN . . . . . . . . . . . . . . . . . . . . . . . . . . . 179
E .5 Configuration Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181
E .6 Ablation Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
E .7 Training and Augmentation time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183
E .8 Augmentation Strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 184
E .9 Graph Distance Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185
E .10 Gaussian Mixture Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 186
E .11 Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 186
E .12 Fisher-Guided GMM Augmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187
E .13 Softmax Confidence and Entropy Distributions . . . . . . . . . . . . . . . . . . . . . . . . . 187
LIST OF FIGURES

Figure 3.1 Result for the spectral clustering task on the Cora graph . . . . . . . . . . . . . . . 36
Figure 4.1 Effect of GCN’s depth on sparse and dense subgraphs . . . . . . . . . . . . . . . . 43
Figure 4.2 Illustration of ADMP-GNN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
Figure 5.1 Robustness guarantees on Cora . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64
Figure 6.1 Illustration of our RobustCRF approach . . . . . . . . . . . . . . . . . . . . . . . . . 70
Figure 7.1 Illustration of GRATIN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
Figure 7.2 The density of the average influence scores . . . . . . . . . . . . . . . . . . . . . . 89
Figure a.1 Dirichlet Energy variation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
Figure a.2 The synthetic graph generated from an SBBA . . . . . . . . . . . . . . . . . . . . . 126
Figure a.3 k-core distributions of three different BA models . . . . . . . . . . . . . . . . . . . . 127
Figure a.4 Result for the spectral clustering task on Cora . . . . . . . . . . . . . . . . . . . . . 127
Figure b.1 The synthetic graphs extracted from Computers and Photo . . . . . . . . . . . . . . 141
Figure c.1 A toy example on how to randomly partition the set [0, 1] . . . . . . . . . . . . . . . 158
Figure c.2 Difference in output for the GCN and our GCORN . . . . . . . . . . . . . . . . . . . 164
Figure d.1 The radius on the lower bound stated in Lemma 6.3.3 . . . . . . . . . . . . . . . . . 169
Figure e.1 Effect of Filtering Augmented Representations on Test Accuracy . . . . . . . . . . . 188
Figure e.2 Softmax Confidence and Entropy Distributions for IMDB-BIN . . . . . . . . . . . . . 188
Figure e.3 Softmax Confidence and Entropy Distributions for IMDB-MUL . . . . . . . . . . . . 189
Figure e.4 Softmax Confidence and Entropy Distributions for MUTAG . . . . . . . . . . . . . . 189
Figure e.5 Softmax Confidence and Entropy Distributions for PROTEINS . . . . . . . . . . . . 189
Figure e.6 Softmax Confidence and Entropy Distributions for DD . . . . . . . . . . . . . . . . . 189

L I S T O F TA B L E S

Table 2.1 Graph Shift Operators. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12


Table 2.2 Statistics of the node classification datasets. . . . . . . . . . . . . . . . . . . . . . . 24
Table 2.3 Statistics of the graph classification datasets. . . . . . . . . . . . . . . . . . . . . . 24
Table 3.1 The result of the spectral clustering task on synthetic graphs . . . . . . . . . . . . . 37
Table 3.2 Performance of CGCN, CGATv2 and other vanilla models . . . . . . . . . . . . . . 38
Table 3.3 Performance of the Local-Global based CGSO . . . . . . . . . . . . . . . . . . . . 39
Table 4.1 Comparison of ADMP-GCN training paradigms ALM and ST . . . . . . . . . . . . . 47
Table 4.2 Comparison of highest accuracy for GCN and ADMP-GCN ST . . . . . . . . . . . . 47
Table 4.3 Classification accuracy for the baselines based on the GCN backbone . . . . . . . 49
Table 4.4 Classification accuracy for the baselines based on the GIN backbone . . . . . . . . 49
Table 5.1 Attacked classification accuracy of the models after the feature attacks . . . . . . . 63
Table 5.2 Attacked classification accuracy of the models after the structural attacks . . . . . . 65
Table 6.1 Attacked classification accuracy of RobustCRF . . . . . . . . . . . . . . . . . . . . 74
Table 6.2 Attacked classification accuracy of RobustCRF when combined with other approaches 75

xiv
L I S T O F TA B L E S xv

Table 6.3 Attacked classification accuracy of RobustCRF with respect to structural attacks . . 76
Table 7.1 Classification accuracy of the data augmentation baselines and GRATINon the GCN
backbone . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
Table 7.2 Classification accuracy of the data augmentation baselines and GRATINon the GIN
backbone . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
Table 7.3 Robustness performance against structure corruption . . . . . . . . . . . . . . . . . 91
Table a.1 Statistics of the node classification datasets used in our experiments. . . . . . . . . 120
Table a.2 Differenet initialization of the weights in CGSO . . . . . . . . . . . . . . . . . . . . . 121
Table a.3 Hyperparameters used in our experiments. . . . . . . . . . . . . . . . . . . . . . . . 122
Table a.4 Performance of CGCN, CGATv2 and other vanilla models on additional datasets . 123
Table a.5 Classification accuracy of CSGC on additional datasets . . . . . . . . . . . . . . . . 124
Table a.6 Classification accuracy of CSGC on additional datasets . . . . . . . . . . . . . . . . 124
Table a.7 Classification accuracy (± standard deviation) of the models on different benchmark
node classification datasets. The higher the accuracy (in %) the better the model. 125
Table a.8 Classification accuracy of CH2GCN . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
Table a.9 Detailed graph properties of the used datasets. . . . . . . . . . . . . . . . . . . . . 135
Table a.10 Graph Properties of the used datasets and the corresponding learned hyperparam-
eters in GAGCN w/ Degree . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136
Table a.11 Graph Properties of the used datasets and the corresponding learned hyperparam-
eters in GAGCN w/ K-Core . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
Table a.12 Graph Properties of the used datasets and the corresponding learned hyperparam-
eters in GAGCN w/ PageRank . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
Table a.13 Graph Properties of the used datasets and the corresponding learned hyperparam-
eters in GAGCN w/ Count of walks. . . . . . . . . . . . . . . . . . . . . . . . . . . . 138
Table b.1 Statistics of the node classification datasets used in our experiments. . . . . . . . . 140
Table b.2 The average time needed for each training setting for different datasets. . . . . . . 141
Table b.3 Hyperparameters used in our experiments. . . . . . . . . . . . . . . . . . . . . . . . 142
Table b.4 Comparison of ADMP-GCN training paradigms ALM and ST . . . . . . . . . . . . . 143
Table b.5 Comparison of highest accuracy for GCN and ADMP-GCN ST . . . . . . . . . . . . 143
Table b.6 Comparison of ADMP-GCN training paradigms ALM and ST . . . . . . . . . . . . . 144
Table b.7 Comparison of ADMP-GIN training paradigms ALM and ST . . . . . . . . . . . . . 144
Table b.8 Comparison of ADMP-GIN training paradigms ALM and ST . . . . . . . . . . . . . 145
Table b.9 Comparison of highest accuracy for GIN and ADMP-GIN ST . . . . . . . . . . . . . 145
Table b.10 Comparison of highest accuracy for GIN and ADMP-GIN ST . . . . . . . . . . . . . 145
Table b.11 Classification accuracy for the baselines based on the GIN backbone . . . . . . . . 146
Table c.1 Attacked classification accuracy after both structural attacks and node feature
attacks application . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155
Table c.2 Performance of GCN and our proposed GCORN model . . . . . . . . . . . . . . . . 156
Table c.3 Mean training time analysis of a our GCORN . . . . . . . . . . . . . . . . . . . . . . 156
Table c.4 Attacked classification accuracy after the attack application . . . . . . . . . . . . . . 160
Table c.5 Classification accuracy before (clean) and after the attack. . . . . . . . . . . . . . . 161
Table c.6 Attacked classification accuracy using GIN after the attack application. . . . . . . . 161
Table c.7 Statistics of the node classification datasets used in our experiments. . . . . . . . . 162
Table c.8 Statistics of the graph classification datasets used in our experiments. . . . . . . . 162
Table d.1 The optimal RobustCRF’s hyperparameters for the feature-based attacks. . . . . . 169
Table d.2 The optimal RobustCRF’s hyperparameters for the structure-based attacks. . . . . 169
Table d.3 Statistics of the node classification datasets used in our experiments. . . . . . . . . 170
xvi L I S T O F TA B L E S

Table d.4 Inference time for different values of the number of iterations/samples. . . . . . . . 171
Table d.5 Attacked classification accuracy on the OGBN-Arxiv dataset. . . . . . . . . . . . . . 171
Table e.1 Classification accuracy for the data augmentation baselines based on the GIN
backbone . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
Table e.2 Ablation study on the density estimation scheme - GCN . . . . . . . . . . . . . . . 183
Table e.3 Ablation study on the density estimation scheme - GIN . . . . . . . . . . . . . . . . 183
Table e.4 Mean training and augmentation time in second . . . . . . . . . . . . . . . . . . . . 184
Table e.5 Classification accuracy for the data augmentation baselines using the GCN backbone184
Table e.6 Classification accuracy for the data augmentation baselines using the GIN backbone185
Table e.7 Statistics of the graph classification datasets used in our experiments. . . . . . . . 187
Table e.8 The optimal number of Gaussian distributions in the GMM . . . . . . . . . . . . . . 187
Part I

INTRODUCTION
INTRODUCTION
1
R aphs provide a natural representation of relationships and interactions between entities, making

G them a fundamental data structure in numerous domains. Formally, a graph consists of a set
of nodes and edges that connect pairs of nodes. Each node can be associated with features,
providing additional context or attributes to the graph structure. Graph-structured data arises in diverse
applications, including social networks, where users and their relationships are modeled as graphs
(Backstrom and Leskovec, 2011), biological networks representing molecular structures (Liu et al., 2017),
knowledge graphs encoding entities and relationships for semantic search (Galkin et al., n.d.), and
transportation networks mapping roads, cities, and routes (Schichl and Sellmann, 2015). Graphs are
ubiquitous, capturing complex and high-dimensional relationships in data that cannot be modeled effectively
with traditional Euclidean structures.
The need to extract meaningful insights from graph-structured data has led to the development of Graph
Representation Learning techniques. These methods aim to derive low-dimensional embeddings that
preserve the structural and feature-based properties of the graph. Unlike traditional machine learning
approaches that operate on feature spaces, graph representation learning leverages the inherent topology
of graphs to capture relationships and dependencies between nodes. Such embeddings are essential
for downstream tasks like node classification, link prediction, and graph classification. These tasks
facilitate applications such as community detection in social networks (Malliaros and Vazirgiannis, 2013a),
recommendation systems (Son and Kim, 2017), and molecular property prediction (Stokes et al., 2020).
To address the challenges of learning on graphs, researchers have developed Graph Neural Networks
(GNNs), which extend traditional neural networks to operate directly on graph structures. GNNs build upon
the principles of graph representation learning by incorporating both node features and graph connectivity
into their computational models. This shift has enabled a more systematic approach to processing graph
data compared to earlier heuristic-based methods. Specifically, GNNs propagate information through
iterative message-passing mechanisms, allowing nodes to aggregate information from their neighbors. This
iterative process mimics how information spreads through networks, making GNNs particularly effective for
tasks requiring relational reasoning.
One prominent class of GNNs is Message Passing Neural Networks (MPNNs) (Kipf and Welling, 2017a;
Xu et al., 2019b), which operate in two stages: a message-passing stage where nodes exchange information
with their neighbors, and an update stage where nodes refine their embeddings based on received
messages. Popular architectures, such as Graph Convolutional Networks (GCNs), leverage convolutional
operators to learn hierarchical features from graph data. These developments have significantly advanced
the ability to learn rich representations from graphs, laying the foundation for more complex machine
learning applications.
Graph Neural Networks have demonstrated remarkable success in both academic and industrial applica-
tions. In social media analysis, platforms like Twitter and LinkedIn use GNNs for recommendation systems
and content filtering (Ying et al., 2018). E-commerce platforms such as Amazon and Alibaba leverage
GNNs for product recommendations (Wang et al., 2019). In the pharmaceutical industry, GNNs aid in
modeling molecular structures and predicting drug interactions, leading to the discovery of new antibiotics
(Stokes et al., 2020).

3
4 INTRODUCTION

1.1 THE CHALLENGES IN GNNS

Despite their success, GNNs face several challenges. One such challenge is over-smoothing, where
repeated message passing leads to the loss of node identity. Scalability is another issue, as large-scale
graphs require efficient processing techniques. Additionally, GNNs are vulnerable to adversarial attacks
that manipulate graph structure or node features. There are also other important challenges, including
robustness, generalization, and improving graph representation learning techniques. Addressing these
challenges opens avenues for improving GNN architectures, making them more expressive, scalable, and
resilient to adversarial perturbations.
— Generalization refers to the ability of a GNN to perform effectively on unseen data or under distribution
shifts that inevitably occur in real-world scenarios.
— Robustness addresses the resilience of GNNs to adversarial perturbations or random noise in both
node features and graph structure, ensuring reliable performance even in malicious or unpredictable
settings.
— Improved representation learning continues to push the boundaries of expressivity and interpretability,
enabling GNNs to capture more relational information and complex structural patterns.

1.2 OV E RV I E W O F C O N T R I B U T I O N S

This thesis explores the theoretical foundations and practical advancements in Graph Neural Networks
(GNNs). By leveraging insights from graph geometry and spectral analysis, this research addresses key
challenges such as robustness and generalization of GNNs, focusing on improving their efficiency and
performance in complex real-world applications.

G R A P H R E P R E S E N TAT I O N L E A R N I N G .
A major focus of my PhD research focuses on adapting
the message-passing scheme in GNNs to align with the varying characteristics of different graphs and
learning tasks. This work has resulted in two works. In one of them, I combined two key concepts, Graph
Shift Operator (GSO) and node centrality criteria, to enhance the performance of GNNs. GSOs, such as
the adjacency and graph Laplacian matrices, are a fundamental concept from graph signal processing,
providing a way to understand how information propagates or “shifts” across a graph. A commonality of the
most frequently used GSOs is their ability to encode purely local information in the graph, with the adjacency
matrix encoding neighborhoods in the graph and the graph Laplacians relying on the node degree, a local
centrality metric, to normalize the adjacency matrix. Building on this, I proposed and studied a new family
of operators called Centrality Graph Shift Operators (CGSOs), which integrate the global position of nodes
within the graph (Abbahaddou et al., 2024c). I performed a comprehensive spectral analysis to explore their
fundamental properties, including eigenvalue structures and expansion behavior, examining how these
operators influence information spread across the graph. In the second work, I addressed the issue that
standard GNNs use a fixed number of message-passing steps for all nodes, ignoring the fact that different
nodes may require different depths of propagation. Through empirical analyses on real-world datasets
and further validated by synthetic experiments, we observed that the optimal number of message-passing
layers varies significantly depending on a node’s characteristics. To tackle this, I developed ADMP-GNN, a
framework that dynamically adjusts the number of message-passing layers for each node. This approach
applies to any model following the message-passing paradigm. Evaluations on node classification tasks
demonstrate that ADMP-GNN outperforms baseline GNNs by tailoring the depth of message passing to
the nodes’ unique needs.
1.2 O V E R V I E W O F C O N T R I B U T I O N S 5

GNN’S ROBUSTNESS. Another significant aspect of my PhD research was devoted to study and
improve GNNs’ robustness against adversarial attacks (Abbahaddou et al., 2024a,b; Ennadir et al., 2024).
Building on theoretical findings, I connected the expected robustness of GNNs to the orthonormality of their
weight matrices and consequently propose an attack-independent and more robust variant of the GCN,
called the Graph Convolutional Orthonormal Robust Networks (GCORNs) (Abbahaddou et al., 2024a). To
do so, I make use of an algorithm first, which we refer to as Bjorck Orthonormalization. In another project,
we introduce NoisyGNNs (Ennadir et al., 2024), a novel defense method that incorporates noise into the
architecture of the underlying model. We establish a theoretical connection between noise injection and the
enhancement of GNN robustness using Renyi Divergence, highlighting the effectiveness of our approach.
And as most existing defense techniques primarily concentrate on the training phase of GNNs, involving
adjustments to message passing architectures or pre-processing methods, I also proposed RobustCRF, a
post-hoc approach aiming to enhance the robustness of GNNs at the inference stage using Conditional
Random Fields (Abbahaddou et al., 2024b).

G N N ’ S G E N E R A L I Z AT I O N .
Another interesting topic in GNN research is generalization. Despite their
impressive capabilities, GNNs face significant challenges related to generalization, particularly when
handling unseen or out-of-distribution (OOD) data (Guo et al., 2024; Li et al., 2022). I introduced a
novel approach, GRATIN, based on Gaussian Mixture Models (GMMs), for graph data augmentation that
enhances both the generalization and robustness of GNNs (Abbahaddou et al., 2025). Using the universal
approximation property of GMMs, we can sample new graph representations to effectively control the upper
bound of Rademacher Complexity, ensuring improved generalization of GNNs.
These contributions collectively advance the theoretical understanding and practical implementation of
GNNs, making them more robust, generalizable, and efficient for diverse applications.
PRELIMINARIES
2
H ischapter provides an introduction to the key concepts and essential background information

T necessary for understanding this thesis. We begin by presenting the notation and fundamental
characteristics of graphs, including core definitions such as nodes, edges, node degree, adjacency
matrices, and random walks, which form the basis for subsequent discussions. Additionally, we explore
graph-based concepts such as Graph Shift Operators (GSOs), which serve as representations of graphs
and play a central role in spectral analysis, as well as node centrality measures, which are used to identify
influential nodes within a graph. Furthermore, we introduce Graph Neural Networks (GNNs) and their
applications, particularly in tasks related to node and graph classification. To evaluate our methods, we
outline a variety of benchmarks encompassing both synthetic and real-world datasets. Collectively, these
elements establish the groundwork for the methodologies and experiments presented in this thesis.

2.1 N O TAT I O N

Let G = (V , E ) be an undirected graph where V is the set of nodes and E is the set of edges. We will
denote by n = |V | and m = |E | the number of nodes and number of edges, respectively. Let N ( j) denote
the set of neighbors of a node j ∈ V , i. e. N ( j) = {i : ( j, i ) ∈ E }. The degree deg(i ) of a node i ∈ V is
equal to its number of neighbors, i. e. equal to |N ( j)| for a node j ∈ V . A graph is commonly represented
by its adjacency matrix A ∈ [0, 1]n×n where the (i, j)-th element ai,j of this matrix is equal to the weight of
the edge between the i-th and j-th node of the graph and a weight of 0 in case the edge does not exist. In
some settings, the nodes of a graph might be annotated with feature vectors. We use X ∈ Rn×K to denote
the node features where K is the feature dimensionality. X is constructed by stacking up the feature vectors
xi ∈ RK for each node i ∈ V . A path between two nodes i, j ∈ V is a sequence of nodes (i0 , i1 , . . . , ik ) such
that i0 = i, ik = j, and (ik′ −1 , ik′ ) ∈ E for all k ′ = 1, . . . , k. A graph is said to be connected if there exists
a path between every pair of nodes in the graph. A graph is said to be connected if there exists a path
between any pair of nodes in the graph. We call a graph non-empty if it has at least one vertex (i.e., n ≥ 1),
and we say it is finite when its set of nodes (and hence edges) is finite in size. An isolated node is a node
with no neighbors, i.e., its degree is zero. We denote by

δ(G) = min{deg( j)} and ∆(G) = max{deg( j)}, (2.1)


j∈V j∈V

the smallest and largest degree of the graph, respectively.

2.2 GRAPH PROPERTIES

Graphs exhibit a wide range of characteristics and can be classified into different types based on their
structural properties. Below, we briefly outline some common characteristics and graph types that are
relevant in various context:

7
8 PRELIMINARIES

2.2.1 Heterogeneity

A heterogeneous graph, also referred to as a multi-relational graph, is defined as G = V , E , TV , TE ,
where TV is the set of node types and T E is the set of edge types. Each node i ∈ V belongs to exactly one
type c ∈ TV . Hence, we can write

V c, V c ∩ V c = ∅ for c ̸= c′ .
[
V= (2.2)
c∈TV

Each edge e ∈ E belongs to exactly one relation type r ∈ T E . Accordingly,



E r, E r ∩ E r = ∅ for r ̸= r ′ .
[
E= (2.3)
r ∈T E

Heterogeneity arises in many complex systems. For example, in knowledge graphs where heterogeneity
arises from the presence of various types of entities and relationships (Li et al., 2024b). Likewise, in
biological networks, heterogeneity emerges from the coexistence of multiple biological entities that interact
through a variety of processes (Yu and Gao, 2022). As a result, one often constructs type-aware adjacency
structures, e. g., separate adjacency matrices for each edge type r ∈ TE .
In contrast, a homogeneous graph consists of a single type of node and edge, making it structurally
simpler and easier to process. For instance, social networks, where nodes typically represent individuals
and edges denote friendships or connections, are often modeled as homogeneous graphs (Yang et al.,
2019). While this simplicity facilitates computational efficiency, homogeneous graphs are inherently limited
in their ability to represent the complexity of relationships.

2.2.2 Homophily

Homophily is the principle that similar nodes tend to connect with each other. In many social and
biological networks, entities sharing an attribute, e. g., labels or biological function, are more likely to form
edges. Graph homophily is measured using the homophily ratio. Formally, let there be a labeling function
ϕ : V → L that assigns each node i ∈ V to a label ϕ(i ) from a label set L. We define the homophily ratio
H as:

{ (i, j) ∈ E : ϕ(i ) = ϕ( j)}


γh = . (2.4)
m
A value of γh close to 1 indicates strong homophily, implying that most edges connect nodes having
the same label. A value of γh close to 0 indicates heterophily, meaning most edges connect nodes with
different labels or attributes.
In tasks such as node classification, strong homophily often makes label propagation or graph-based
semi-supervised learning more effective, since neighbors are likely to share labels (Luan et al., 2024;
Platonov et al., 2024). This intuition is well illustrated by the widely used Label Propagation (LP) algorithm,
which is commonly applied to node classification and serves as a simple yet effective baseline for Graph
Neural Networks (GNNs) (Wang and Leskovec, 2020). Specifically, labels are iteratively spread through
the graph during label propagation, allowing unlabeled nodes to inherit the labels of their neighbors. When
homophily is strong, the likelihood that neighboring nodes share the same label is high, making the
propagation process more accurate and reliable.
2.3 G R A P H E X P L O R AT O N : W A L K S A N D R A N D O M W A L K S 9

2.2.3 Sparsity

n ( n −1)
A graph with n nodes has a maximum of 2 edges if it is an undirected, homogeneous and simple
graph, i. e., with no self-loops. The graph G is said to be sparse if the actual number of edges m = |E | is
n ( n −1)
much smaller than this maximum. In other words, m ≪ 2 . One way to quantify sparsity is via the
edge density, defined as
2m
ρ= . (2.5)
n ( n − 1)
Sparsity significantly influences the computational complexity of graph algorithms. For example, many
methods in graph signal processing, or graph neural networks, can be optimized when the graph is sparse.
For example, sparse representations reduce the complexity of matrix-vector multiplications by processing
only existing edges, avoiding redundant calculations over non-existent connections.

2.2.4 Connectivity

A graph is connected if every pair of nodes is linked by some path in G . More precisely, G is connected if
for any two nodes u, v ∈ V there exists a sequence of edges that starts at u and ends at v. Formally:

∀i, j ∈ V ∃k ∈ N ∃i0 , . . . , ik ∈ V , (i0 = i, i1 ), (i1 , i2 ), . . . , (ik−1 , ik = j) ∈ E . (2.6)


If G is not connected, it can be partitioned into connected components, where each connected component
C ⊆ V is a maximal set of nodes such that any two nodes within C are connected by a path lying entirely in
C.
Connectivity in graphs has a high impact on GNN performances. On one hand, well-connected graphs
enable features from labeled nodes to propagate effectively to unlabeled nodes. On the other hand, in dis-
connected graphs or graphs with isolated nodes, information exchange is limited to connected components,
preventing the model from capturing global patterns and leading to fragmented representations that hinder
learning and generalization (Abbahaddou, Lutzeyer, and Vazirgiannis, 2023).

By accounting for heterogeneity, homophily, sparsity, and connectivity, one can better model and un-
derstand the diverse structures and behaviors observed in real-world graphs. Each of these properties
influences the choice of algorithms and models, from conventional graph mining techniques to mod-
ern graph neural networks. In the next section, we will delve into walks and random walks, which are
fundamental concepts in graph machine learning.

2.3 G R A P H E X P L O R AT O N : W A L K S A N D R A N D O M W A L K S

In this section, we introduce fundamental concepts related to graph explorations, focusing on both
deterministic and random walks. Walks provide a straightforward means of moving through the graph, while
random walks incorporate probability distributions at each step, offering a powerful framework for studying
and simulating various graph-based phenomena. We begin by defining and discussing walks, and then
proceed to formalize random walks as Markov processes on graphs.
10 PRELIMINARIES

2.3.1 Walks

As defined in Definition 2.3.1, a walk in a graph is a sequence of nodes where each consecutive pair is
connected by an edge. While walks can revisit nodes, a path is a specific type of walk where all vertices
are distinct (except possibly the start and end nodes in a cycle), ensuring no repetitions.
Definition 2.3.1 (Walk on a Finite Graph). Let G = (V , E ) be an graph with node set V and edge set
E . A walk of length T in G is a finite sequence of vertices (i0 , i1 , . . . , it ) such that ii , ii+1 ∈ E for every
t = 0, 1, . . . , T − 1. If i0 = i T , the walk is said to be closed.
The number of walks of length t from a vertex i to a vertex j is precisely given by the (i, j)-th entry of
(A + I)t , where A is the adjacency matrix of the graph. Normalized walks are a variant of walks that
account for differences in node degrees to balance contributions across the graph. For normalized walks,
one typically uses the Normalised Adjacency matrix  = D−1/2 (A + I)D−1/2 ; in these cases, the entries
of Ât give normalized number of walk from i to j in exactly t steps.
Building on this concept of a walk, we can introduce a probability distribution at each step to define a
random walk. This stochastic perspective allows us to model processes where each successive move
is chosen randomly, providing a powerful framework for studying and simulating diverse phenomena on
graphs. Next, we introduce the formal definition of random walks.

2.3.2 Random Walks

Random walks form the foundation of many graph-based algorithms and analyses. They provide a
probabilistic framework to model through a graph, where each step is determined by the connectivity of
nodes. Random walks are particularly useful for studying graph properties, evaluating node importance,
and designing algorithms for tasks such as node ranking, clustering, and semi-supervised learning (Chien
et al., 2020; Malliaros and Vazirgiannis, 2013a; Yang et al., 2019). We give the detailed mathematical
definition of Random Walks in Definition 2.3.2.
Definition 2.3.2 (Random Walk on a Finite Graph). Let G = (V , E ) be a finite, undirected graph, and let
δG : V × V −→ N denote the shortest-path distance between any two vertices, i. e., ρ(i, j) is the minimum
number of edges in a path connecting i and j. A random walk on G is a sequence of random variables
( It )t≥0 taking values in V , such that:
1. Markov Property: For all t ≥ 0 and all vertices i0 , i1 , . . . , it , k ∈ V satisfying

≤ 1 for 0 ≤ t′ ≤ t − 1,

ρ i t ′ , i t ′ +1 (2.7)
we have

P It+1 = k ( It , . . . , I0 ) = (it , . . . , i0 ) = P It+1 = k It = it .


 
(2.8)
In other words, the probability of moving to k at time t + 1 depends only on the current vertex it , not
on the earlier history.
2. Transition Probabilities: For each vertex vt , the probability of transitioning to a vertex z at time t + 1
is given by

0, if δG (it , k) > 1,


P It+1 = k It = it

= ait ,k
(2.9)
, if δG (it , k) ≤ 1.


deg(it )

2.4 G R A P H S H I F T O P E R AT O R S 11

In other words, if It is at vertex i, then It+1 is determined by moving to an adjacent vertex k, using a
randomly and uniformly chosen edge connecting j to k, independent of the past history of the walk. This
includes the possibility that It+1 = j, which may happen if there is a loop at i.
The distribution of the initial step I0 of the walk is called the initial distribution, characterized by the
probabilities P( I0 = i ) for i ∈ V. If I0 = i0 almost surely, the random walk is said to start from i0 . The
process ( It ) forms a Markov chain with state space V and transition matrix P = p j,k j,k∈V given by,

0, if δG (v, z) > 1,
p j,k = (2.10)
 a j,k , if dΓ ( j, k ) ≤ 1,
deg( j)

where a(v, z) represents the adjacency matrix coefficient.


A random walk on G admits a unique stationary distribution π. For an undirected graph, it is well known
that

deg( j)
π ( j) = , (2.11)
∑i∈V deg(i )
meaning that in the long run, the probability of being at a node v is proportional to its degree (Kowalski,
2019; Lovász, 1993). This stationary behavior is essential for deriving analytical results that connect
long-term graph exploration with the structural properties of the graph. It plays a key role in various
applications, including graph-based ranking, community detection, and beyond (Page, 1999).

2.4 G R A P H S H I F T O P E R AT O R S

In this section, we introduce Graph Shift Operators (GSOs), a class of matrices that plays a central role
in processing and analyzing graph-structured data. By encoding a graph’s connectivity through sparse,
edge-based relationships, GSOs facilitate both theoretical insights into spectral properties and practical
algorithms for tasks such as clustering, community detection, and spectral analysis. In particular, degree-
normalized GSOs have received significant attention for their robustness and interpretability in applications
ranging from signal processing on graphs to graph neural networks.

2.4.1 Definition of Graph Shift Operators

Understanding the structure and behavior of graph-structured data often requires efficient and insightful
analysis tools. Graph Shift Operators (GSOs) play an important role in the analysis of graph-structured data.
Degree-normalized GSOs, such as the Random-walk Normalized Laplacian (Modell and Rubin-Delanchy,
2021), have been widely employed in spectral analysis and signal processing on graphs. These GSOs
have many properties allowing great insight in the connectivity of nodes. Formally, a graph G = (V , E ), can
be represented by an adjacency matrix A = [ ai,j ] ∈ R N × N where ai,j = 1 if ei,j ∈ E and ai,j = 0 otherwise.
Analyzing the spectra of the adjacency matrix provides information about the basic topological properties
of the underlying graphs (Cvetkovic, Doob, and Sachs, 1980). For example, according to the Gerschgorin
theorem, every eigenvalue of an adjacency matrix A lies in at least one of the circular discs with center aii
and radii ∑i≤ j | aij |. Furthermore, the largest eigenvalue of A is an upper bound of the average degree and
a lower bound of the largest degree (Cvetković, Rowlinson, and Simić, 2009; Sarkar and Jalan, 2018).
In addition to the spectral properties of the adjacency matrix, there are alternative graph representations
that provide deep insights into the topology of the underlying graph. One often-used representation is the
symmetrically normalized Laplacian matrix defined by Lsym = I − D−1/2 AD−1/2 , where D ∈ Rn×n is the
12 PRELIMINARIES

Table 2.1 – Graph Shift Operators.

Graph Shift Operator Description when V = D

A Adjacency matrix
L = D−A Unnormalised Laplacian matrix
Q = D+A Signless Laplacian matrix
Lrw = I − D−1 A Random-walk Normalised Laplacian matrix
Lsym = I − D−1/2 AD−1/2 Symmetric Normalised Laplacian matrix
 = D−1/2 (A + I)D−1/2 Normalised Adjacency matrix
H = D −1 A Mean Aggregation matrix

degree matrix, i.e., a diagonal matrix defined as Dii = ∑in=1 aij . There are other graph representations with
particularly interesting spectral properties, such as the random-walk Normalised Laplacian (Modell and
Rubin-Delanchy, 2021) and the Signless Laplacian matrices (Cvetković and Simić, 2009). All these graph
representations belong to the family of Graph Shift Operators (GSOs). Examples of such matrices appear
in Table 2.1. Note that the adjacency matrix A, the unnormalized Laplacian L, and other variations listed
there all satisfy the structural criteria of GSOs, differing mainly by how they weight each edge.

Definition 2.4.1. Given an arbitrary graph G = (V , E ), a Graph Shift Operator S ∈ R N × N is a matrix


satisfying Sij = 0 for i ̸= j and (i, j) ∈
/ E (Mateos et al., 2019) and Sij ̸= 0 for i ̸= j and (i, j) ∈ E .

In addition to the classical or fixed GSOs, parametrized GSOs are learned during the optimization
process. These parametrized operators are a fundamental component of many modern GNN architectures
and allow the model to adapt and capture complex patterns and relationships in the graph data. For
example, the work of PGSO (Dasoulas, Lutzeyer, and Vazirgiannis, 2021a) parametrizes the space of
commonly used GSOs leading to a learnable GSO that adapts to the dataset and learning task at hand.

2.4.2 Spectral Clustering

One important application of GSOs is spectral clustering (Von Luxburg, 2007), a technique for partitioning
a graph into distinct clusters (or communities), leveraging the spectral properties of GSOs to partition
graphs into meaningful clusters. The process begins by computing GSO’s eigenvectors. Nodes are then
embedded into a lower-dimensional spectral space using the eigenvectors corresponding to the smallest
non-zero eigenvalues. A clustering algorithm, such as K-means, is applied in this spectral space to group
nodes, and the resulting clusters are mapped back to the original graph, revealing its community structure.
Spectral clustering is widely used in practice. For instance, in social networks, it can be employed to detect
user communities sharing similar interests or social ties (Wahl and Sheppard, 2015). In collaborative
filtering and recommender systems, spectral clustering may help identify user or item groups with similar
2.4 G R A P H S H I F T O P E R AT O R S 13

Algorithm 1: Spectral Clustering using the Centrality GSOs


1 Inputs: Graph G , A GSO Φ, Number of clusters to retrieve C.
1. Compute the eigenvalues {λ}in=1 and eigenvectors {u}in=1 of Φ;
2. Consider only the eigenvectors U ∈ R N ×C corresponding to the C largest eigenvalues;
3. Cluster rows of n, corresponding to nodes in the graph, using the K-Means algorithm to retrieve a node
partition P with C clusters;
P = K-Means(U, C )
return P ;

interaction patterns, providing a basis for improved recommendation mechanisms (Hoseini, Hashemi, and
Hamzeh, 2012). In Algorithm 1, we outline the steps of the Spectral Clustering Algorithm using CGSOs.

2.4.3 Spectral Properties of GSOs

Beyond clustering, studying the spectral properties of GSOs gives rise to important graph expansion
inequalities. For example, the Cheeger constant measures the bottleneck of a graph, and Cheeger’s
inequality quantitatively links this bottleneck to the spectrum of the graph Laplacian. These relationships
are crucial for understanding how well-connected a graph is. The relevant definitions and results are
summarized below.
Definition 2.4.2 (Cheeger constant). Let G = (V , E ) be a finite graph.
1. For any disjoint subsets of vertices V1 , V2 ⊂ V , we denote by E (V1 , V2 ) the set of edges of G with
one extremity in V1 and one extremity in V2 ,

E (V1 , V2 ) = {(i, j) ∈ E | {i, j} ∩ V1 ̸= ∅, {i, j} ∩ V2 ̸= ∅}. (2.12)


and we denote by E (V1 ) the set E (V1 , V \ V1 ) of edges with one extremity in V1 , and one outside V1 .
2. The Cheeger constant hC (G) is defined by
|E (W )|
 
1
hC (G) = min W ⊂ V, W ̸= ∅, |W | ≤ |V | , (2.13)
|W | 2
with the convention that hC (G) = +∞ if G has at most one vertex.
Proposition 2.4.3 (Discrete Cheeger inequality, (Cheeger, 1970a)). Let G = (V , E ) be a connected,
non-empty, finite graph without isolated vertices. We have
 
2∆(G)
1 − ϱr ( H ) ≤ λ1 (Lrw ) ≤ hC (G) (3.32)
δ(G)2
where δ(G) and ∆(G) represent the minimum and maximum degree of the graph, λ1 (Lrw ) is the normalized
spectral gap, i. [Link] smallest non-zero eigenvalue of Lrw , and ϱr ( H ) is Equidistribution radius of G , i. [Link]
maximum of the absolute values |λ| for λ an eigenvalue of the Mean Aggregation matrix H = D−1 A which
is different from ±1.
Proposition 2.4.4 (Discrete Buser inequality, (Buser, 1982)). Let G = (V , E ) be a connected, non-empty,
finite graph without isolated nodes. We have
q
hC (G) ≤ ∆(G) 2λ1 (Lrw ), (2.14)
14 PRELIMINARIES

where ∆(G) represents the maximum degree of the graph, and λ1 (Lrw ) is the Normalized Spectral Gap of
G , i. [Link] smallest non-zero eigenvalue of Lrw .

Proposition 2.4.5 (Spectral Properties of the Unnormalised Laplacian Matrix L). Let G = (V , E ) be a finite
graph. The Unnormalized Laplacian matrix L of G exhibits the following key spectral properties:
1. The multiplicity of the 0 eigenvalue is equal to the number of connected components in G and the
corresponding eigenvectors are indicator vectors establishing which vertex is an element of which
connected component (Von Luxburg, 2007).
2. Let G be a connected graph with diameter Diam(G). Then, the Unnormalised Laplacian matrix L of
G has at least Diam(G) + 1 distinct eigenvalues (Brouwer and Haemers, 2011).

The inequalities in Propositions 2.4.3, 2.4.4 and 2.4.5 demonstrate how the graph’s connectivity and
expansion properties are intimately tied to the spectrum of the Random-walk Normalised Laplacian and the
Unnormalised Laplacian matrices. These inequalities underscore the importance of careful selection and
design of GSOs in tasks such as clustering, community detection, and analysis of graphs. For example,
the celebrated Cheeger’s Inequality establishes a bound on the edge expansion of a graph via its spectrum
(Cheeger, 1970b). We recall the concept of the Cheeger constant, which quantifies how well-connected a
graph is by measuring the ratio between the number of edges leaving a set of vertices and the size of that
set. The Cheeger Inequality in Proposition 2.4.3 shows how the smallest non-zero eigenvalue of Lrw is
bounded above and below by terms involving the Cheeger constant. By relating combinatorial quantities
(e. g., the Cheeger constant) to spectral properties (e. g., eigenvalues of the Laplacian), we gain powerful
tools for assessing and controlling how graph-structured data behaves under transformations induced by
GSOs.
Overall, these theoretical underpinnings highlight why Graph Shift Operators are central to modern
approaches for learning, analyzing, and processing data on graphs. By leveraging their spectral properties,
researchers can design algorithms that better respect the structure of real-world networks, whether they
are social, biological, or technological in nature.

2.5 NODE CENTRALITY

Node centrality measures the relative importance of a node within a graph based on its connectivity to
other nodes. Node centrality has numerous real-world applications, including identifying influential nodes in
social networks (Malliaros, Rossi, and Vazirgiannis, 2016), and ranking webpages in search engines (Brin
and Page, 1998). It is also used in recommendation systems (Son and Kim, 2017), and traffic flow analysis
(Borgatti, 2005), where understanding the central nodes helps optimize performance and decision-making.
In addition to degree centrality, which provides a local measure based on a node’s direct connections,
various other centrality metrics have been proposed to capture different aspects of a node’s importance. In
this section, we explore several centrality measures, including the k-core, PageRank, and Walk Count, each
providing unique insights into the importance of nodes within a graph. These metrics serve as valuable
tools for understanding the structural properties of nodes in complex networks. Below, we introduce three
centrality measures that are used in this manuscript.

2.5.1 k-core

We leverage the k-core decomposition of a graph, which captures how well-connected nodes are within
their neighborhood (Malliaros et al., 2020). The process of k-core decomposition involves iteratively
2.5 N O D E C E N T R A L I T Y 15

Algorithm 2: k-core Decomposition


1 Inputs: Non-Empty Graph G with a diameter larger than 1.
1.
2. Initialization: Set G0 = G and k = 1.
3. Iterative Pruning: At each iteration, construct the subgraph Gk by removing all vertices in Vk−1 with degree
less than k:
Gk = Gk−1 \ {i ∈ Vk−1 | degGk−1 (i ) < k} (2.15)

where degGk−1 (i ) denotes the degree of vertex i in the subgraph Gk−1 .


4. Termination: Repeat the pruning process until Gk is either empty or no vertices with degree less than k
remain. Increment k at the end of each iteration:

k = k+1

return P ;

removing vertices with degree less than k until no such vertices remain. Mathematically, the k-core can be
constructed using Algorithm 2.
The core number core(i ) of a vertex i is the highest k for which i is part of the k-core. Formally:

core(i ) = max{k | i ∈ Gk }. (2.16)

This core number reflects the robustness of a node’s connectivity within the graph. Nodes with higher core
numbers are considered more central as they reside in denser subgraphs, indicating strong interconnections
with their neighbors.
The k-core decomposition provides insights into the hierarchical structure of the graph, revealing densely
connected regions and identifying influential nodes based on their participation in these cores. It is
particularly useful in applications such as community detection, and network robustness (Dey et al., 2020;
Peng, Kolda, and Pinar, 2014).

2.5.2 PageRank

The PageRank algorithm, originally developed for ranking webpages, evaluates the importance of nodes
within a connected web graph by assigning a numerical weight to each node. This weight reflects the
node’s relative significance, determined by the structure of incoming and outgoing links (Brin and Page,
1998). The concept has since been generalized to apply to any graph structure. A node’s PageRank
score represents the probability of a random walk visiting that node, making it a fundamental metric for
measuring node importance in various networks, particularly in web search algorithms. Mathematically, the
PageRank algorithm operates by modeling a random walk over the graph. If the graph G contains nodes
with no outgoing links, the random walk is no longer well-defined. To address this, a natural approach
involves allowing the random walk to jump uniformly at random to any node in V with a probability of n1 .
The transition matrix Pi,j is then defined as,

aij

deg+ ( j)
, if deg+ ( j) > 0,
Pij = (2.17)
1
n, otherwise.
16 PRELIMINARIES

Here, deg+ ( j) is the out-degree of node j, i. e., the total number of edges directed away from node j.
While this ensures that the random walk is well defined, the process may still lack irreducibility, as certain
nodes may remain inaccessible from others.
To enforce irreducibility, a damping factor α ∈ [0, 1] is introduced. With probability α, the random walk
follows the transition matrix P, and with probability 1 − α, it restarts at a random node. The modified
transition matrix P(α) becomes,
1
P(α) = αP + (1 − α) 11T , (2.18)
n
where 1 is the column vector of ones. The stationary distribution π (α) , known as the PageRank vector,
satisfies the equation,
1
π (α) = απ (α) P + (1 − α) 1T . (2.19)
n
This stationary distribution π (α) gives the PageRank score for each node, indicating the steady-state
probability of visiting that node during a random walk (Hollocou, Bonald, and Lelarge, 2016).

2.5.3 Walk Count

The Walk Count centrality quantifies the importance of a node based on the number of walks of a given
length ℓ starting from that node. This measure captures higher-order connectivity patterns within the
graph, going beyond direct neighbors to account for paths that traverse intermediate nodes. Formally, let
A ∈ Rn×n denote the adjacency matrix of the graph G , and 1 ∈ Rn represent the vector of ones. The
vector Pℓ-paths ∈ Rn records the number of paths of length ℓ originating from each node i, defined as,

Pℓ-paths = Aℓ 1, (2.20)

where Aℓ represents the ℓ-th power of the adjacency matrix, corresponding to the number of walks of length
ℓ between nodes. For the special case of ℓ = 2, Pℓ-paths can be further expressed as, diag(P2-paths ) =
W M13 1D, where W M13 is the motif adjacency matrix introduced by (Benson, Gleich, and Leskovec, 2016),
counting the number of open bidirectional wedges (motif M13 ), D is the degree matrix of the graph, and
diag(·) is an operator that converts a vector into a diagonal matrix, placing the vector’s elements on the
diagonal while setting all off-diagonal elements to zero. This motif network captures higher-order structures
and gives new insights into the organization of complex systems.

2.6 GRAPH NEURAL NETWORKS

Before the advent of GNNs, machine learning on graphs relied on methods such as handcrafted
embedding techniques. Approaches like Node2Vec (Grover and Leskovec, 2016), DeepWalk (Perozzi,
Al-Rfou, and Skiena, 2014), and LINE (Tang et al., 2015) used random walks or other strategies to
generate low-dimensional embeddings of nodes. While these methods achieved reasonable success,
they lacked the ability to generalize to unseen graphs or directly leverage node and edge attributes during
training. Furthermore, their reliance on static embeddings made them suboptimal for capturing dynamic
and multi-hop dependencies within a graph.
GNNs address these limitations by extending the deep learning paradigm to graph-structured data.
Unlike traditional approaches, GNNs can dynamically learn representations of nodes, edges, and even
entire graphs by iteratively aggregating and transforming information from local neighborhoods. This
capability enables GNNs to capture both the structural and semantic properties of graphs, making them
highly effective for a wide range of tasks such as node classification, link prediction, and graph classification.
2.6 G R A P H N E U R A L N E T W O R K S 17

In the following sections, we delve deeper into the key architectures of GNNs, focusing on their design
principles and mathematical foundations. We start with the basic framework of the message-passing
scheme, a fundamental concept in most recent GNN models.

2.6.1 Key GNN Architectures

A Graph Neural Network (GNN) consists of multiple neighborhood aggregation layers, where each GNN
layer generates new node representations relying on the graph structure G = (V , E ) and the nodes’ feature
vectors from the previous layer. Suppose we have a GNN model that contains L neighborhood aggregation
(0)
layers. Let also hi denote the initial feature vector of node i, i. e. hi = xi . At each layer (ℓ > 0), the
(ℓ)
hidden state h j of a node j is updated as follows:
 
(ℓ) (ℓ−1)
aj = AGGREGATE(ℓ) hi : i ∈ N ( j) , (2.21)
 
(ℓ) (ℓ−1) (ℓ)
h j = COMBINE(ℓ) h j , aj , (2.22)

where AGGREGATE(·) is a permutation invariant function that maps the feature vectors of the neighbors of a
node v to an aggregated vector. This aggregated vector is passed along with the previous representation
(ℓ−1)
of j, i. e. h j , to the COMBINE(·) function which combines those two vectors and produces the new
representation of v. In the Graph Convolutional Network (GCN) (Kipf and Welling, 2017a), the AGGREGATE
function aggregates features from the neighbors of each node using a normalized sum, and the COMBINE
function is a linear transformation followed by a non-linear activation as follows:
1

(ℓ) (ℓ−1)
aj = p hi (2.23)
i ∈N ( j)∪{ j} deg(i )deg( j)
 
(ℓ) (ℓ)
hj = σ W (ℓ a j , (2.24)

where W (ℓ) ∈ Rdℓ−1 ,dℓ are learnable weight matrix, dℓ is the dimension of the hidden representation at the
ℓ-th layer, and σ is a non-linear activation function, often chosen as ReLU.
In the Graph Attention Network (GAT) (Veličković et al., 2018), the AGGREGATE function incorporates a
learnable attention mechanism, which assigns different importance scores to neighbors. The attention
coefficients are computed as follows:
  
(ℓ−1) (ℓ−1)
exp LeakyReLU a⊤ [W (ℓ) hi ∥ W (ℓ) h j ]
(ℓ)
αij =    , (2.25)
(ℓ−1) (ℓ−1)
∑k∈N ( j)∪{ j} exp LeakyReLU a⊤ [W (ℓ) hk ∥ W (ℓ) h j ]
′ ′
where a ∈ R2dℓ is a learnable vector, W (ℓ) ∈ Rdℓ ×dℓ is a learnable weight matrix, and ∥ denotes the
vector concatenation. The new nodes representation are expressed as follows:


(ℓ) (ℓ) (ℓ−1)
aj = αij W (ℓ) hi , (2.26)
i ∈N ( j)
 
(ℓ) (ℓ)
hj = σ aj . (2.27)

GATv2 (Brody, Alon, and Yahav, 2022) improves upon GAT by introducing a more expressive attention
mechanism. Instead of applying the linear attention coefficients directly, GATv2 computes attention scores
as:
18 PRELIMINARIES

 
(ℓ)  (ℓ−1) (ℓ−1) 
eij = a⊤ σ W (ℓ) hi ∥ hj , (2.28)
(ℓ)
where eij is the unnormalized attention score. These scores are normalized using a softmax function:
 
(ℓ)
exp eij
(ℓ)
αij =  . (2.29)
(ℓ)
∑k∈N ( j)∪{ j} exp ekj

The rest of the aggregation and combination steps are similar to GAT.
Graph Isomorphism Network (GIN) (Xu et al., 2019b) takes a different approach by focusing on maximiz-
ing the expressive power of the aggregation function. The AGGREGATE function in GIN uses a sum, which is
theoretically proven to be as powerful as the Weisfeiler-Lehman graph isomorphism test. It is defined as:


(ℓ) (ℓ−1)
aj = hi , (2.30)
i ∈N ( j)∪{ j}

and the combination function is:


 
(ℓ) (ℓ−1) (ℓ)
hj = MLP(ℓ) (1 + ϵ(ℓ) )h j + aj , (2.31)

where ϵ(ℓ) is a learnable parameter, and MLP(ℓ) is a multi-layer perceptron that applies non-linear transfor-
mations.
These architectures highlight the diverse strategies in designing AGGREGATE and COMBINE functions
to achieve expressive and effective graph learning. After L layers of message passing, each node’s
representation incorporates information from its -hop neighborhood. Formally, an L-hop neighborhood of a
node includes all nodes that are at most edges away from . This property allows GNNs to capture local
and higher-order structural information from the graph. However, the effective receptive field of a GNN is
inherently limited by the number of layers, which can impact its ability to model long-range dependencies
in graphs. Furthermore, as increases, challenges such as over-smoothing, where node representations
become indistinguishable, may arise. Understanding this aggregation behavior is critical for designing
effective GNN architectures tailored to specific tasks and datasets.
After L iterations of neighborhood aggregation, to produce a graph-level representation, GNNs apply a
permutation invariant readout function, e. g. the sum operator, to nodes feature as follows,
 
( L)
hG = READOUT h j : j ∈ V . (2.32)

2.6.2 Challenges in GNNs

Despite their growing popularity and success across various applications, GNNs encounter several chal-
lenges that limit their performance. Specifically, we delve into issues such as oversmoothing, oversquashing,
robustness, generalization, and scalability.

[Link] Oversmoothing

mai In graph neural networks (GNNs), oversmoothing refers to the phenomenon where, as the number
of layers increases, the node representations become indistinguishable from each other. This typically
happens because repeated message passing leads to information from different nodes blending together
excessively, making it difficult to distinguish nodes based on their features. This phenomenon can be
2.6 G R A P H N E U R A L N E T W O R K S 19

theoretically understood through the graph Laplacian and quantified using the Dirichlet Energy (Zhou et al.,
2021), defined as

(ℓ) 2
(ℓ) hj
n o 1 hi
= ∑ aij p
(ℓ)
Edir hi |i∈V −p ,
2 i,j 1 + deg(i ) 1 + deg( j)
2
n o
(ℓ)
where hi | i ∈ V represents the node representations at the ℓ-th GNN layer. If the Dirichlet Energy, it
indicates that representations of connected nodes become close to one another. Consequently, all node
features collapse to an almost constant signal on the graph, making it impossible to distinguish between
individual nodes based on their learned embeddings.

[Link] Oversquashing

Oversquashing refers to the compression of information from a large neighborhood into a fixed-size
vector, making it difficult for message-passing neural networks (MPNNs) to propagate and preserve all
relevant information across the graph. This problem becomes especially severe in tasks involving long-
range dependencies, where distant nodes must exchange signals through multiple hops (Dwivedi et al.,
2022). Addressing these challenges requires careful architectural design, such as limiting the number of
layers, incorporating skip connections, or using alternative aggregation mechanisms to better preserve and
propagate information across the graph.

[Link] Robustness

Robustness in GNNs refers to their ability to maintain performance under various perturbations, such as
noise, adversarial attacks, or incomplete data. In the context of adversarial attacks, referred to as adversarial
robustness, the focus is on reducing the effect of changes to degrade model performance. These attacks
typically involve perturbing node features or graph structure and can be categorized into evasion attacks,
which occur during testing, and poisoning attacks, which target the training process (Zügner, Akbarnejad,
and Günnemann, 2018). Enhancing robustness against such attacks can be achieved through methods
like adversarial training, which incorporates adversarial examples during training to bolster resistance, or
certifiable robustness, which involves developing models with formal guarantees against specific types of
perturbations (Bojchevski and Günnemann, 2019a).

2.6.3 Generalization

GNNs have demonstrated significant success in tasks such as node and graph classification, yet they
often face difficulties in generalizing, particularly to unseen or out-of-distribution (OOD) (Li et al., 2022; Tang
and Liu, 2023). These challenges are intensified when training data is limited in size or diversity. Various
theoretical tools, including Vapnik-Chervonenkis dimension (Pfaff et al., 2020), Rademacher complexity
(Yin, Kannan, and Bartlett, 2019), and algorithm stability (Pfaff et al., 2020), have been employed to provide
generalization bounds for GNNs. Notably, (Liao, Urtasun, and Zemel, 2020) was the first to establish
generalization bounds for GCNs and message-passing neural networks using the PAC-Bayesian approach,
while Neural Tangent Kernels have also been utilized to analyze the generalization of infinitely wide GNNs
trained via gradient descent (Du et al., 2019; Jacot, Gabriel, and Hongler, 2018).
20 PRELIMINARIES

2.6.4 Expressivity

Expressivity in GNNs measures their capacity to model and distinguish complex graph structures and
capture rich node or edge-level dependencies. This is influenced by both the underlying architecture
and theoretical constraints. The expressivity of GNNs is often analyzed through the Weisfeiler-Lehman
(WL) test of graph isomorphism, which iteratively updates node labels based on their neighbors and
assesses the ability to distinguish non-isomorphic graphs (Xu et al., 2019b). Standard message-passing
GNNs are constrained by the expressive power of the 1-WL test (Li and Leskovec, 2022). To address
this, researchers have proposed enhancements such as higher-order GNNs, which leverage higher-order
WL tests to model interactions among larger substructures (Morris et al., 2019). Additional methods
include incorporating subgraph information to capture higher-order features (Bevilacqua et al., 2022) and
using attention mechanisms to enable more nuanced feature aggregation by assigning different levels of
importance to neighbors (Joshi et al., 2023).

[Link] Scalability

Scalability is a critical challenge for GNNs, particularly when applied to large-scale graphs with millions or
even billions of nodes and edges (Hu et al., 2020a). Training a GNN on the large graphs is computationally
expensive, necessitating methods like sampling or clustering to handle such scale while maintaining
performance. Sampling-based methods reduce the computational cost by sampling a fixed number of
neighbors at each layer rather than processing the full neighborhood (Hamilton, Ying, and Leskovec, 2017).
Cluster-based methods partition the graph into smaller subgraphs or clusters, allowing GNNs to process
these independently and then combine the results. Additionally, efficient architectures, such as scalable
attention mechanisms and sparse matrix operations, further optimize performance (Gao et al., 2022).

2.7 G R A P H L E A R N I N G TA S K S

Now, we introduce a set of important tasks performed on graph data,

2.7.1 Semi-Supervised Node classification

In semi-supervised node classification, the goal is to predict the labels of nodes within a graph when
only a subset of nodes has known labels. Formally, we are given a set of labels Ytrain = {yi , i ∈ Vtrain } for
a fixed subset of nodes Vtrain ⊂ V . The task is to leverage both the labeled nodes in Vtrain and the graph
structure G = (V , E ) to learn a model that predicts the labels Ytest of the remaining test nodes Vtest , the
unlabeled set.
Semi-supervised node classification is commonly encountered in applications like social networks (e. g.,
predicting user attributes), citation networks (e. g., predicting article categories), and molecular graphs (e. g.,
predicting functional groups) (Shchur et al., 2018a; Wu et al., 2018). The GNN architecture is well-suited
for this task since it uses the structure and feature information of neighboring nodes, effectively propagating
label information from labeled nodes through the graph to inform predictions on unlabeled nodes.

2.7.2 Graph Classification

In the Graph Classification task, the goal is to predict a label for each graph in a given set G . Formally,
we are provided with a dataset G = {G1 , G2 , . . . , G p }, where each graph Gi is defined by its structure
(Vi , Ei ), the set of nodes and edges, and its node attributes XGi . Each graph Gi is associated with a label
2.8 G R A P H D I S TA N C E M E T R I C S 21

yi ∈ Y , where Y is the label space, such as binary or multi-class categories. The task is to learn a function
f : G → Y , e. g., a GNN, that maps graphs to their respective labels. During training, the model learns
from a labeled training set Gtrain to extract meaningful patterns. The learned model is then evaluated on a
separate test set Gtest , to assess its ability to generalize and correctly classify new graphs.
Graph classification is widely used in many real-word applications. For example, GNNs can predict
molecular properties where molecules are represented as graphs with atoms as nodes and chemical
bonds as edges (Godwin et al., 2022). This enables efficient screening of potential drug candidates and
understanding molecular interactions. Similarly, in Bioinformatics, graph classification plays a critical role in
analyzing biological networks, such as protein-protein interaction networks or gene regulatory networks, to
identify disease related patterns or classify functional behaviors of biological entities (Hsu et al., 2022).

2.8 G R A P H D I S TA N C E M E T R I C S

Graphs are widely used to model complex systems across various domains, including social networks,
biological systems, and transportation networks. Comparing graphs is a fundamental task that enables
the analysis of structural and functional similarities or differences. Graph distance metrics provide a
mathematical framework to quantify these similarities and differences by evaluating the adjacency structures
and node attributes. These metrics are particularly useful in applications such as graph classification,
clustering, and anomaly detection. This section introduces key graph distance measures, focusing on both
structural and feature-based differences, along with methods to account for node alignment issues.
There are several approaches to compute graph distance metrics, such as Edit Distance, which measures
the number of edit operations (additions, deletions, or substitutions) required to transform one graph into
another (Gao et al., 2010); Graph Kernels, which map graphs into a feature space where distances can be
computed using inner products (e.g., the Weisfeiler-Lehman kernel and Random Walk kernels) (Nikolentzos
and Vazirgiannis, 2020; Shervashidze et al., 2011); Spectral Distances, which compare eigenvalues and
eigenvectors of adjacency or Laplacian matrices to capture structural differences (Jovanović and Stanić,
2012); Optimal Transport based Distances, which measures distances by solving optimization problems
for node alignments and weights—particularly useful when graphs differ in node ordering (Chen et al.,
2020a); and Embedding-Based Methods, which represent graphs in vector spaces using graph embeddings
and compute Euclidean distances between these representations. By focusing on both structural and
feature-based differences and accounting for node alignment issues, these measures provide a robust
foundation for numerous graph analysis tasks.
Building upon the overview of key graph distance measures, we now formalize these metrics within a
mathematical framework. This involves defining graph distances that capture structural and feature-based
changes. Let us consider the graph space (A, ∥·∥A ) and the feature space (X, ∥·∥X ), where ∥·∥A and
∥·∥X denote the norms applied to the graph structure and features, respectively. When considering only
structural changes, with fixed node features, the distance between two graphs G1 , G2 is defined as

∥G1 − G2 ∥ = ∥A1 − A2 ∥A , (2.33)


where A1 , A2 are respectively the adjacency matrices of G1 , G2 ,. The norm ∥·∥G can be expressed using
different metrics. The Hamming Distance is defined as:
∥A1 − A2 ∥ H = ∑ 1(A1 (i, j) ̸= A2 (i, j)), (2.34)
i,j

which counts the number of differing entries between adjacency matrices. The Frobenius Norm is given by:
s
∥A1 − A2 ∥ F = ∑(A1 (i, j) − A2 (i, j))2 , (2.35)
i,j
22 PRELIMINARIES

measuring the element-wise differences. Finally, the Spectral Norm is expressed as:
∥A1 − A2 ∥2 = σmax (A1 − A2 ), (2.36)
where σmax denotes the largest singular value of the matrix difference, capturing global structural deviations.
These norms provide different perspectives on measuring graph similarity depending on the application. If
both structural and feature changes are considered, the distance extends to:

∥G1 − G2 ∥ = α∥A1 − A2 ∥G + β∥ X1 − X2 ∥X , (2.37)


where X1 , X2 are the node feature matrices of G1 , G2 respectively, and α, β are positive hyperparameters
controlling the contribution of structural and feature differences.
In cases where the node alignment between the two graphs is unknown, we must take into account node
permutations. The distance between the two graphs is then defined as
 
∥G1 − G2 ∥ = min α∥A1 − PA2 PT ∥G + β∥ X1 − PX2 ∥X , (2.38)
P∈Π
where Π is the set of permutation matrices. A permutation matrix P is a binary matrix where each row
and column contains exactly one entry equal to 1, and all other entries are 0. Mathematically, it satisfies
PP T = P T P = I, where I is the identity matrix. The set Π includes all such matrices of a given size,
representing all possible node reorderings.
Permutation matrices are necessary because graphs often do not have a fixed node ordering. For
instance, two structurally identical graphs may have different node indices. Direct comparison without
alignment may lead to misleading results. The permutation matrix P reorders nodes in G2 to align with G1 ,
ensuring structural and feature comparisons are valid.
Using Optimal Transport, we find the minimum distance over the set of permutation matrices, correspond-
ing to the optimal matching between nodes in the two graphs. This approach generalizes graph distance
computation to handle cases where node correspondences are unknown.

2.9 SYNTHETIC GRAPHS

Synthetic graphs play a crucial role in the study of graph-based algorithms models (Tsitsulin et al.,
2022), and machine learning approaches. These graphs are artificially generated using predefined rules
and statistical models, enabling researchers to evaluate the scalability, robustness, and effectiveness of
various algorithms under controlled conditions. Unlike real-world graphs, synthetic graphs provide flexibility
in tuning parameters such as size, density, degree distribution, and clustering, making them ideal for
hypothesis testing and benchmarking. In this section, we will introduce two well-known types of synthetic
graphs, the Stochastic Block Model (SBM) and the Barabási–Albert (BA) model,which are employed in this
thesis analyze graph properties.

2.9.1 Stochastic Block Model (SBM)

The Stochastic Block Model (SBM) is a generative model designed to create graphs with community
structures. It partitions nodes into K distinct blocks or communities and specifies intra- and inter-community
edge probabilities.
Formally, let G = (V , E ) represent a graph with n nodes. Nodes are divided into K communities, and
each node i belongs to a community ci ∈ {1, . . . , K }. The probability of an edge between two nodes i and j
depends on their community memberships:
P( aij = 1) = pci ,c j , (2.39)
2.10 B E N C H M A R K S A N D D ATA S E T S 23

where pci ,c j is the probability of an edge existing between communities ci and c j . In a common special
case, edges within the same community share a fixed probability values p and edges between different
communities share a fixed probability values q. When p ≫ q, the network exhibits pronounced homophily,
i. e., nodes within the same community form densely connected subgroups that are relatively sparse
between groups. Conversely, if p ≈ q, the graph becomes more uniform, i. e., the model becomes similar to
an Erdős–Rényi random graph, making it more challenging to detect distinct communities. SBM is widely
used for studying community detection algorithms and modeling social and biological networks (Lee and
Wilkinson, 2019; Wilkinson, 2018).

2.9.2 Barabási–Albert (BA) Model

The Barabási–Albert (BA) model generates scale-free networks characterized by a power-law degree
distribution, mimicking real-world graphs such as social and biological systems. This model builds a graph
through a preferential attachment mechanism, which simulates the rich-get-richer phenomenon, i. e., nodes
with higher degrees are more likely to attract new connections. The construction process involves the
following steps,
1. Start with a small set of m0 nodes.
2. At each step, add a new node with m ≤ m0 edges.
3. Connect the new node to existing nodes with a probability proportional to their degree:

deg(i )
P( i ) = , (2.40)
∑ j deg( j)

where deg(i ) is the degree of node i.


Despite its simplicity, the BA model captures fundamental features of complex graphs, offering a
foundation for understanding networks that evolve over time. For example, in social networks, new users
tend to connect preferentially with individuals who already have a large number of connections, leading to
the formation of hubs that represent highly influential users.

2.10 B E N C H M A R K S A N D D ATA S E T S

In this section, we present the datasets and benchmarks used to evaluate the models in this thesis.
The experiments focus on two key tasks, node and graph classifications. For each task, we describe the
datasets, their key characteristics Tables 2.2 and 2.3 provide detailed statistical summaries of each dataset.

2.10.1 Datasets for the Node Classification Task

In this thesis, we run experiments on the node classification task using the citation networks Cora,
CiteSeer, and PubMed (Sen et al., 2008), the co-authorship networks CS and Physiscs (Shchur et al.,
2018b), the citation network between Computer Science arXiv papers OGBN-Arxiv (Hu et al., 2020a), the
Amazon Computers and Amazon Photo networks (Shchur et al., 2018b), the non-homophilous datasets
Penn94 (Traud, Mucha, and Porter, 2012), genius (Lim and Benson, 2021), deezer-europe (Rozemberczki
and Sarkar, 2020) and arxiv-year (Hu et al., 2020a), and the disassortative datasets Chameleon, Squirrel
(Rozemberczki, Allen, and Sarkar, 2021), and Cornell, Texas, Wisconsin from the WebKB dataset (Lim
et al., 2021a). For the Cora, CiteSeer, and Pubmed datasets, we used the provided train/validation/test
splits. For the remaining datasets, we followed the framework in (Lim et al., 2021a; Rozemberczki, Allen,
24 PRELIMINARIES

and Sarkar, 2021). Characteristics and information about the datasets utilized in the node classification
part of the study are presented in Table 2.2.

Table 2.2 – Statistics of the node classification datasets.

Dataset #Features #Nodes #Edges #Classes Edge Homophily


Cora 1,433 2,708 5,208 7 0.809
CiteSeer 3,703 3,327 4,552 6 0.735
PubMed 500 19,717 44,338 3 0.802
CS 6,805 18,333 81,894 15 0.808
arxiv-year 128 169,343 1,157,799 5 0.218
chameleon 2,325 2,277 62,792 5 0.231
Cornell 1,703 183 557 5 0.132
deezer-europe 31,241 28,281 185,504 2 0.525
squirrel 2,089 5,201 396,846 5 0.222
Wisconsin 1,703 251 916 5 0.206
Texas 1,703 183 574 5 0.111
Photo 745 7,650 238,162 8 0.827
ogbn-arxiv 128 169,343 2,315,598 40 0.654
Computers 767 13752 491,722 10 0.777
Physics 8,415 34,493 495,924 5 0.931
Penn94 4,814 41,554 2,724,458 3 0.470

2.10.2 Datasets for the Graph Classification Task

For the graph classification task, we evaluate our models on widely used datasets from the GNN
literature, specifically IMDB-BINARY, IMDB-MULTI, PROTEINS, MUTAG, and DD, all sourced from the
TUD Benchmark (Ivanov, Sviridov, and Burnaev, 2019). These datasets consist of either molecular or
social graphs. Detailed statistics for each dataset are provided in Table 2.3.

Table 2.3 – Statistics of the graph classification datasets.

Dataset #Graphs Avg. Nodes Avg. Edges #Classes

IMDB-BINARY 1,000 19.77 96.53 2


IMDB-MULTI 1,500 13.00 65.94 3
MUTAG 188 17.93 19.79 2
PROTEINS 1,113 39.06 72.82 2
DD 1,178 284.32 715.66 2
2.11 E VA L U AT I O N M E T R I C S 25

2.11 E VA L U AT I O N M E T R I C S

In this section, we describe the evaluation metrics used to assess the performance of the proposed
models for both node and graph classification tasks.

2.11.1 Node Classification

Node classification aims to predict labels for nodes in a graph based on their features and connectivity.
Given a graph G = (V , E ) with |V | nodes and a set of true labels Y = {yi | yi ∈ {1, . . . , C }} for each node
i ∈ V, where C is the number of classes, we define the predicted labels as Yb = {ybi }.
The classification accuracy for node classification is computed as:
1
|Vtest | i∈∑
Accuracy = 1(yi = ybi ), (2.41)
Vtest

where Vtest denotes the set of test nodes, and 1(·) is the indicator function that equals 1 if its argument is
true and 0 otherwise.

2.11.2 Graph Classification

For graph classification, the task is to assign a label yi ∈ {1, . . . , C } to each graph Gi in a set of graphs
G = {G1 , G2 , . . . , G N }, where N is the number of graphs.
Let ybi denote the predicted label for graph Gi . The classification accuracy for graph classification is then
defined as:
1 Ntest
Ntest i∑
Accuracy = 1(yi = ybi ), (2.42)
=1
where Ntest is the number of test graphs, and 1(·) is the indicator function.

2.12 S O F T WA R E A N D TO O L S

In this section, we outline the software libraries, frameworks, and tools used to conduct the experiments
and analyses presented in this thesis. The selected tools were chosen based on their efficiency, scalability,
and support for graph-based computations, making them particularly suitable for graph neural network
(GNN) research.

2.12.1 Programming Languages and Frameworks

— Python: The primary programming language used for implementing models, algorithms, and experi-
ments. Its extensive ecosystem supports scientific computing and machine learning tasks.
— PyTorch: (Paszke et al., 2019) A flexible and efficient deep learning framework utilized for imple-
menting and training neural network architectures, including GNN models.

2.12.2 Graph Processing Libraries

— NetworkX: (Hagberg, Swart, and Schult, 2008) Employed for graph creation, manipulation, and
analysis. It provides a comprehensive collection of graph algorithms and visualization tools, aiding in
initial data exploration and preprocessing.
26 PRELIMINARIES

— PyTorch Geometric (PyG): (Fey and Lenssen, 2019) A specialized library for deep learning on
graphs, enabling the implementation of graph neural networks. It supports tasks such as node
classification, graph classification, and link prediction through optimized message-passing operations.
Part II

R E P R E S E N TAT I O N L E A R N I N G I N G R A P H N E U R A L N E T W O R K S
R E T H I N K I N G G R A P H S H I F T O P E R AT O R S I N G N N S F O R G R A P H
R E P R E S E N TAT I O N L E A R N I N G
3
Shift Operators (GSOs), such as the adjacency and graph Laplacian matrices, play a

G
RAPH
fundamental role in graph theory and graph representation learning. Traditional GSOs are typically
constructed by normalizing the adjacency matrix by the degree matrix, a local centrality metric. In
this work, we instead propose and study Centrality GSOs (CGSOs), which normalize adjacency matrices
by global centrality metrics such as the PageRank, k-core or count of fixed length walks. We study spectral
properties of the CGSOs, allowing us to get an understanding of their action on graph signals. We confirm
this understanding by defining and running the spectral clustering algorithm based on different CGSOs on
several synthetic and real-world datasets. We furthermore outline how our CGSO can act as the message
passing operator in any Graph Neural Network and in particular demonstrate strong performance of a
variant of the Graph Convolutional Network and Graph Attention Network using our CGSOs on several
real-world benchmark datasets.

3.1 INTRODUCTION

We propose and study a new family of operators defined on graphs that we call Centrality Graph Shift
Operators (CGSOs). To insert these into the rich history of matrices representing graphs and centrality
metrics, the two concepts married in CGSOs, we begin by recalling major advances in these two topics
in turn (readers interested purely in recent developments in Graph Representation Learning and Graph
Neural Networks are recommended to begin reading in Paragraph 3 of this section). The study of graph
theory and with it the use of matrices to represent graphs have a long-standing history. Graph theory is
often said to have its origins in 1736 when Leonard Euler posed and solved the Königsberg bridge problem
(Euler, 1736). His solution did not involve any matrix calculus. In fact, it seems that the first matrix defined
to represent graph structures is the incidence matrix defined by Henri Poincaré in 1900 (Poincaré, 1900). It
is difficult to pinpoint the first definition of adjacency matrices, but by 1936 when the first book on the topic
of graph theory was published by Dénes König adjacency matrices had certainly been defined and began
to be used to solve graph theoretic problems (König, 1936). Two seemingly concurrent works in 1973
defined an additional matrix structure to represent graphs that later became known as the unnormalized
graph Laplacian (Donath and Hoffman, 1973; Fiedler, 1973). Then, it was Fan Chung in her book “Spectral
Graph Theory” published in 1997 who extensively characterized the spectral properties of normalized
Laplacians (Chung, 1997). In the emerging field of Graph Signal Processing (GSP) (Ortega et al., 2018;
Sandryhaila and Moura, 2013) these different graph representation matrices were all defined to belong to a
more general family of operators defined on graphs, the Graph Shift Operators (GSOs). GSOs currently
play a crucial role in graph representation learning research, since the choice of GSO, used to represent a
graph structure, corresponds to the choice of message passing function in the currently much-used Graph
Neural Network (GNN) models.
In parallel to advances in graph representation via matrices, centrality metrics have proved to be insightful
in the study of graphs. Chief among them is the success of the PageRank centrality criterion revealing the
significance of certain webpages (Brin and Page, 1998) and playing a role in the formation of what is now
one of the largest companies worldwide. But also an even older metric, the k-core centrality (Malliaros et al.,
2020; Seidman, 1983), as well as the degree centrality, closeness centrality, and betweenness centrality,

29
30 R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

have proven to be impactful in revealing key structural properties of graphs (Freeman, 1977; Zhang and
Luo, 2017).
A commonality of the most frequently used GSOs is their property to encode purely local information in
the graph, with the adjacency matrix encoding neighborhoods in the graph and the graph Laplacians relying
on the node degree, a local centrality metric, to normalize the adjacency matrix. In this work, we study
a novel class of GSOs, the Centrality GSOs (CGSOs) that arise from the normalization of the adjacency
matrix by centrality metrics such as the PageRank, k-core and the count of fixed length walks emanating
from a given node. Our CGSOs introduce global information into the graph representation without altering
the connectivity pattern encoded in the original GSO and therefore, maintain the sparsity of the adjacency
matrix. We provide several theorems characterizing the spectral properties of our CGSOs. We confirm
the intuition gained from our theoretical study by running the spectral clustering algorithm on the basis of
our CGSOs on 1) synthetic graphs that are generated from a stochastic blockmodel in which each block
is sampled from the Barrabasi-Albert model and 2) the real-world Cora graph in which we aim to recover
the partition provided by the k-core number of each node. We will furthermore describe how our CGSOs
can be inserted as the message passing operator into any GNN and observe strong performance of the
resulting GNNs on real-world benchmark datasets.
In particular, our contributions can be summarized as follows.
(i) We define Centrality GSOs, a novel class of GSOs based on the normalization of the adjacency
matrix with different centrality metrics, such as the degree, PageRank score, k-core number, and the
count of walks of a fixed length,
(ii) We conduct a comprehensive spectral analysis to unveil the fundamental properties of the CGSOs.
Our gained understanding of the benefits of CGSOs is confirmed by running the spectral clustering
algorithm using our CGSOs on synthetic and real-world graphs.
(iii) We incorporate the proposed CGSOs within GNNs and evaluate performance of a Graph Convolu-
tional Network and Graph Attention Network v2 with a CGSO message passing operator on several
real-world datasets.

3.2 R E L AT E D W O R K

This section builds upon the earlier introduction to GNNs by providing a rigorous introduction of their
matrix-based formulation.
Matrix Formulations of Graph Neural Networks. Graph Neural Networks (GNNs) are neural networks
that operate on graph-structured data that is defined as the combination of a graph G = (V , E ), and a node
feature matrix X ∈ Rn×K , containing the node feature vector of node i in its ith row. GNNs are formed by
stacking several computational layers, each of which produces a hidden representation for each node in the
(ℓ)
graph, denoted by H(ℓ) = [hv ]v∈V . A GNN layer ℓ updates node representations relying on the structure
of the graph and the output of the previous layer H(ℓ−1) . Conventionally, the node features are used as
input to the first layer H0 = X. The most popular framework of GNNs is that of Message Passing Neural
Networks (Hamilton, 2020), where the computations are split into two main steps:
Message Passing: Given a node v, this step applies a permutation-invariant function to its neighbors,
denoted by N (v), to generate the aggregated representation,

M(ℓ+1) = Φ(A)H(ℓ) , (3.1)

where Φ(A) : R N × N → R N × N , a function of the adjacency matrix, is the chosen GSO.


3.3 C G S O : C E N T R A L I T Y G R A P H S H I F T O P E R AT O R S 31

Update: In this step, we combine the aggregated hidden states with the previous hidden representation
of the central node v, usually by making use of a learnable function,
H(ℓ+1) = σ (M(ℓ+1) W (ℓ) ), (3.2)
where W (ℓ) ∈ Rdℓ−1 ,dℓ are learnable weight matrices and dℓ is the dimension of the hidden representation
at the ℓ-th layer.
With the emergence and increasing popularity of GNNs, the importance of GSOs has significantly
increased. Numerous GNN architectures, such as notably Graph Convolutional Networks (GCNs), rely
on these operators in their message passing step. In the context of GCNs (Kipf and Welling, 2017b), the
used message passing operator, i.e., the chosen GSO, corresponds to Φ(A) = D1−1/2 AD1−1/2 , where
D1 = D + I is the degree matrix of the graph corresponding to the adjacency matrix A1 = A + I. For Graph
(ℓ)
Attention Networks v2 (GATv2) (Brody, Alon, and Yahav, 2022), (3.1) becomes M(ℓ+1) = Φ(AGATv2 )H(ℓ) ,
(ℓ)
where, in this setting, Φ corresponds to the identity function and the rows of AGATv2 contain the edge-wise
attention coefficients.
As we will present shortly, in this work, we generalize the concept of GSOs to encompass global structural
information beyond node degree. The proposed CGSO framework encapsulates several global centrality
criteria, demonstrating intriguing spectral properties. We further leverage CGSOs to formulate a new class
of message passing operators for GNNs, enhancing model flexibility.
Global Information in GNNs. Besides our GNNs, which leverage the CGSO to make global information
accessible to any given GNN layer, there exists a plethora of other approaches to achieve this goal. These
include for example the PPNP and APPNP (Gasteiger, Bojchevski, and Günnemann, 2019), as well as the
PPRGo (Bojchevski et al., 2020) models that use the PageRank centrality to define a completely new graph
over which to perform message passing in GNNs. The work of Vela et al. (n.d.) extends these models to
consider both the PageRank and k-core centrality. In addition, there is the AdaGCN (Sun, Zhu, and Lin,
2019) and the VPN model (Jin et al., 2021) which propose to message pass using powers of the adjacency
matrix to incorporate global information and increase the robustness of GNNs, respectively. Lee et al.
(2019) propose the Motif Convolutional Networks, that define motif adjacency matrices and then use these
in the message passing scheme. Also the k-hop GNNs of Nikolentzos, Dasoulas, and Vazirgiannis, 2020
consider neighbors several hops away from a given central node in the message passing scheme of a single
GNN layer to consider more global information in a GNN. Additionally there exists a rich and long-standing
literature on spectral GNNs that facilitate global information exchange by explicitly or approximately making
use of the spectral decomposition of the GSO chosen to be the GNN’s message passing operator (Bruna
et al., 2014; Defferrard, Bresson, and Vandergheynst, 2016; Koke and Cremers, 2024). Finally, there is an
arm of research investigating graph transformers, where usually the graph structure is only used to provide
structural encodings of nodes and the optimal message passing operator is learning using an attention
mechanism (Kreuzer et al., 2021; Ma et al., 2023; Rampášek et al., 2022a). All these approaches increase
the computational complexity of the GNN, whereas our CGSO based GNNs maintain the complexity of the
underlying GNN model by preserving the sparsity of the original adjacency matrix.

3.3 C G S O : C E N T R A L I T Y G R A P H S H I F T O P E R AT O R S

In this section, we introduce the Centrality GSOs (CGSO), a family of shift operators that incorporate
the global position of nodes in a graph. We discuss different instances of CGSOs corresponding to widely
used centrality criteria. We further conduct a comprehensive spectral analysis to unveil the fundamental
properties of CGSOs, including the eigenvalue structure and the expansion properties, examining how
these operators influence information spread across the graph. Then, we leverage CGSOs in the design of
flexible GNN architectures.
32 R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

3.3.1 Mathematical Formulation

For a given node i ∈ V , let v(i ) denote a centrality metric associated with i, such as the node degree,
k-core number, PageRank, or the count of walks of specific length starting from node i. The Hilbert space
L2 ( G ) is characterized by the set of functions φ defined on V such that ∑i∈V v(i )| φ(i )| converges, equipped
with the inner product: ⟨ φ1 , φ2 ⟩G = ∑i∈V v(i ) φ1 (i ) φ̄2 (i ). The Markov Averaging Operator on L2 ( G ) is
defined as the linear map MG : φ 7→ MG φ such that
  1
(MG φ) (i ) = V−1 Aφ (i ) =
v (i ) ∑ φ( j),
j∈Ni

where V = diag (v(1), . . . , v( N )) and Ni is the neighborhood set of node i. The form of this Markov
Averaging Operator gives rise to the simplest formulation of our CGSOs, which is a left normalization of the
adjacency matrix by a diagonal matrix containing node centralities on the diagonal, i.e., V−1 A. Note that
the mean aggregation operator, as discussed in Xu et al. (2019b), represents a specific instance of these
CGSOs where the degree corresponds to the chosen centrality metric, namely V = D. We will further
extend the concept of CGSOs in (3.4) where we extend and parameterize these CGSOs. In this paper, we
focus on three global centrality metrics, in addition to the local node degree. We recall the definitions of
these global centrality metrics now.
k-core. The k-core number of a node can be determined in the process of the k-core decomposition of a
graph, which captures how well-connected nodes are within their neighborhood (Malliaros et al., 2020).
The process of k-core decomposition involves iteratively removing vertices with degree less than k until no
such vertices remain. The core number k of a node is then equal to the largest k for which the considered
nodde is still present in the graph’s k-core decomposition. We define Vcore ∈ R N × N to be the diagonal
matrix indicating the core number of each node, i. e.∀i ∈ V , Vcore [i, i ] = core(i ).
PageRank. We choose VPR ∈ R N × N such that, ∀i ∈ V , VPR [i, i ] = (1 − PR (i ))−1 , where PR (i )
corresponds to the PageRank score (Brin and Page, 1998). The PageRank score quantifies the likelihood
of a random walk visiting a particular node, serving as a fundamental metric for evaluating node significance
in various networks.
Walk Count. Here, we consider Vℓ-walks ∈ R N × N , the diagonal matrix indicating the number of walks of
length ℓ starting from each node i, i.e., ∀i ∈ V , Vℓ-walks [i, i ] = Aℓ 1 [i ], where 1 ∈ R N is the vector of ones.


When ℓ = 2, Vℓ-walks corresponds to WM13 1 − D, where WM13 the graph operator presented by Benson,
Gleich, and Leskovec (2016), which corresponds to the count of open bidirectional wedges, i.e., the motif
M13 . This motif network captures higher-order structures and gives new insights into the organization of
complex systems.
In what follows, we delve into the theoretical properties of Markov Averaging Operators, since all three
CGSOs Vcore , VPR and Vℓ-walks are instances of Markov Averaging Operators.

Proposition 3.3.1. The following properties of operator MG hold.


(1) MG is self-adjoint.
(2) MG is diagonalizable in an orthonormal basis,
 itseigenvalues are real numbers, and all eigenvalues
v (i )
have absolute values at most γ = mini∈V deg(i) .

The proof of Proposition 3.3.1 and all subsequent theoretical results in this section can be found in
Appendix a.10. Hence, we have shown in Proposition 3.3.1 that all CGSOs have a real set of eigenvalues,
which is of real use in practice.
In the now following Proposition 3.3.2 we provide the mean and standard deviation of the spectrum of
MG , i.e., the set of MG ’s eigenvalues.
3.3 C G S O : C E N T R A L I T Y G R A P H S H I F T O P E R AT O R S 33

Proposition 3.3.2. The following properties hold for the spectrum of MG .


(1) In a graph G = (V , E ) with multiple connected components C ⊂ V , where each connected com-
ponent C induces a subgraph of G denoted by GC , a complete set of eigenvectors of MG can be
constructed from the eigenvectors of the different MGC , where eigenvectors of MGC are extended to
have dimension N via the addition of zero entries in all entries corresponding to nodes not in the
currently considered component C .
(2) The mean µ(MG ) and standard deviation σ (MG ) of MG ’s spectrum have the following analytic form

1 n 1
µ (MG ) = n ∑ i =1 v ( i ) ,
h
1 1
 2 i1/2
σ (MG ) = n ∑(i,j)∈E v (i ) v ( j )
− µ spϕ .

We define the normalized spectral gap λ1 ( G ) as the smallest non-zero eigenvalue of I − MG . In


Proposition 3.3.4, we link λ1 ( G ) to the expansion properties of the graph. In the literature, we characterize
graph expansion via the expansion or Cheeger constant (Chung, 1997), which measures the minimum ratio
between the size of a vertex set and the minimum degree of its vertices, reflecting the graph’s connectivity.
In our work, we generalize this definition to any centrality metric.

Definition 3.3.3. For a graph G = (V , E ) we define the centrality-based Cheeger constant hv ( G ) as follows

|∂U |
 
1
hv (G) = min | U ⊂ V, |U |v ≤ |V |v , (3.3)
|U | v 2

where |∂U | equals the number of vertices that are connected to a vertex in U but are not in U, and
| · |v : U ⊂ V 7→ ∑i∈V v(i). When the chosen centrality is the degree, hv ( G ) corresponds to the classical
Cheeger constant.

Definition 3.3.3 allows us to establish a link between the spectrum of our considered Markov operators,
i.e., CGSOs, and the centrality-based Cheeger constant in Proposition 3.3.4.

Proposition 3.3.4. Let G be a connected, non-empty, finite graph without isolated vertices. We have,

v2+
 
λ1 ( G ) ≤ 2N hv (G),
v−

where we denote v− = mini∈V v(i ) and v+ = maxi∈V v(i ).

3.3.2 CGNN: Centrality Graph Neural Network

CGSOs, as defined above, normalize the adjacency matrix based on the centrality of the nodes, thereby
providing a refined representation of graph connectivity. Here, we leverage CGSOs to design flexible
message passing operators in GNNs. Incorporating CGSOs within GNNs aims to harness structural
information, enhancing the model’s ability to discern subtle topological patterns for prediction tasks. To
achieve this, we integrate these operators, without loss of generality, in Graph Convolutional Networks
(GCNs) (Kipf and Welling, 2017b) and Graph Attention Networks v2 (GATv2)(Brody, Alon, and Yahav,
2022). We replace the initial shift operator Φ(A) in (3.1), with the proposed CGSOs Φ(A, V), incorporating
different types centrality operators V defined in Section 3.3.1.
It has been shown that the maximum PageRank score converges to zero when the total number of
nodes is very high (Cai et al., 2023), which is the case in many real-world dense graph data (Leskovec
et al., 2010; Leskovec and Mcauley, 2012). Also, the number of walks is high when the expansion of the
34 R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

graph is high. Thus, training a GNN with the proposed CGSOs can lead to numerical instabilities such as
vanishing and exploding gradients. To avoid such issues, we can control the range of the eigenvalues of
CGSOs. We particularly consider a learnable parameterized CGSO framework which is a generalization of
the work of Dasoulas, Lutzeyer, and Vazirgiannis, 2021b. This has the further advantage that the CGSOs
are fit to the given datasets and learning tasks, which leads to more accurate and higher performing graph
representation. The exact formula of the new parametrized CGSO is

Φ(A, V) = m1 Ve1 + m2 Ve2 A a Ve3 + m3 I N , (3.4)

where A a = A + aI N , and (m1 , m2 , m3 , e1 , e2 , e3 , a) are scalar parameters that are learnable via backprop-
agation. Here m1 controls the additive centrality normalization of the adjacency matrix. The parameter e1
controls whether the additive centrality normalization is performed with an emphasis on large centrality
values (for large positive values of e1 ) or with an emphasis on small centrality values (for large negative
values of e1 ). Similarly, we have e2 and e3 controlling the emphasis on large or small centralities, as well as
whether the multiplicative centrality normalization of the adjacency matrix is performed symmetrically or
predominantly as a column or row normalization. The parameter m2 controls the magnitude and sign of the
adjacency matrix term; in particular, a negative m2 corresponds to a more Laplacian-like CGSO, while a
positive m2 gives rise to a more adjacency-like CGSO. Finally, a determines the weight of the self-loops
that are added to the adjacency matrix, and m3 controls a further diagonal regularization term of the CGSO.
More details on the experimental setup are provided in Section 3.5.
In our experiments, we notice the best centrality to vary across datasets, although the walk-based
centrality CGSO appears to be frequently outperformed by the k-core and PageRank CGSO. More
particularly, in some cases e.g. PubMed, it is desirable to use local centrality metrics such as the degree,
while for other datasets e.g., Cornell, it’s preferable to normalize the adjacency with global centrality metrics.
In light of this uncertainty, we can opt for a dynamic, trainable choice of centrality by including both local
and global centrality-based CGSO in our CGNN; this can be done by summing the CGSO of the degree
matrix with the CGSO of a global centrality metric, e.g.,

Φ = Φ(A, D) + Φ(A, Vcore )

.
The parameters m1 , m2 , m3 controlling the magnitude of both the local and global CGSOs are then able
to learn the relative importance of the local and the global CGSO. In Section 3.5, we provide experimental
results for GNNs with such combined CGSOs.
Time Complexity. We recall that the main complexity of our CGCN model is concentrated around
the pre-computation of each centrality score. Computing the degree of all nodes in a graph has a time
complexity of O(|V | + |E |), where |V | is the number of nodes and |E | is the number of edges in the graph
(Cormen et al., 2022). For the PageRank algorithm, each iteration requires one vector-matrix multiplication,
which on average requires O(|V |2 ) time complexity. To compute the core numbers of nodes, we iteratively
remove nodes with a degree less than a specified value until all remaining nodes have a degree greater
than or equal to that value. This operation can be done with a complexity of O(|V | + |E |). Finally, counting
the number of walks of length ℓ for all the nodes can be done via matrix multiplication Aℓ 1 where 1 ∈ R N is
the vector of ones. Since our CGSOs preserve the sparsity pattern of the original adjacency matrix, the
complexity of the GNNs in which the CGSOs are inserted is unaltered.
3.4 A S P E C T R A L C L U S T E R I N G P E R S P E C T I V E O F C G S O S 35

3.4 A SPECTRAL CLUSTERING PERSPECTIVE OF CGSOS

In this section, we analyze CGSOs through the lens of spectral clustering (Ng, Jordan, and Weiss, 2001;
Von Luxburg, 2007). Spectral clustering is a powerful technique that relies on the spectrum of GSOs to
reveal underlying structures within graphs, providing insights into their connectivity properties.

3.4.1 Spectral Clustering on Stochastic Block Barabási–Albert Models

Here, we investigate the behavior of CGSOs in the spectral clustering task on synthetic data. Specifically,
we propose a new graph generator that is a trivial combination of the well-known Stochastic Block Models
(SBM) (Holland, Laskey, and Leinhardt, 1983) and Barabási–Albert (BA) models (Albert and Barabasi,
2002), we call this generator the Stochastic Block Barabási–Albert Models (SBBAM). We will now discuss
the properties and parameterizations of these two graph generators in turn to then discuss their combination
in the SBBAMs.
SBMs. Firstly, in SBMs the node set of the graph is partitioned into a set of K disjoint blocks B1 , . . . , BK ,
where both the number and size of these blocks is a parameter of the model. In SBMs edges are drawn
uniformly at random with probability pij for i, j ∈ {1, . . . , K } between nodes in blocks Bi and B j . Note that
this parameterization is often simplified by the following constraints pij = q if i = j and pij = p if i ̸= j.
SBMs produce graphs which exhibit cluster structure if p ̸= q, which makes them a common benchmark for
clustering algorithms and subject to extensive theoretical study (Abbe, 2018). Note that SBMs can produce
both homophilic graphs if p < q and heterophilic graphs if q > p (Lutzeyer, 2020, Figure 1.2).
BA. The second ingredient of our SBBAMs are the Barabási–Albert (BA) models (Albert and Barabasi,
2002). This model generates random scale-free networks using a preferential attachment mechanism,
which is why these models are also sometimes referred to as preferential attachment (PA) models. In this
PA mechanism we start out with a seed graph and then add nodes to it one-by-one at successive time
steps. For each added node r edges are sampled between the added node and nodes existing in the
graph, where the probability of connecting to existing nodes is proportional to their degree in the graph.
Hence, high degree nodes are more likely to have their degree rise even further than low degree nodes in
future time steps of the generation process (an effect, that is some time referred to as ‘the rich get richer’).
BA models characterize several real-world networks (Barabási and Albert, 1999). A key characteristic of
a BA model is their degree distribution. In Lemma 3.4.1, we prove that the density and connectivity of a
BA model strongly depend on and positively correlate with the hyperparameter r. Thus, we can generate
structurally different BA models by choosing different values of r. Lemma 3.4.1 is proved in Appendix a.11.
Lemma 3.4.1. Let GBA be a Barabási–Albert graph of N nodes generated with the hyperparameters
N0 < N the initial number of nodes, r0 ≤ N02 the initial number of random edges and r the number of added
edges at each time step. Then, the average degree in the network is,
r0 r
deg(G BA ) = 2r + 2 − 2N0 ,
N N
and thus, as the number of nodes grows, i.e., N → ∞, the average degree becomes deg(G BA ) ∼ 2r.
SBBAMs. In our SBBAMs we combine SBMs and BA models, by sampling K BA graphs each of size
|B1 |, . . . , |BK | and with parameters r1 , . . . , rK . We then randomly draw edges between nodes in different BA
graphs, Bi and B j , uniformly at random with probability pij for i, j ∈ {1, . . . , K }. In other words, SBBAMs
trivally extend SBMs to graph in which each block is generated using a BA model. This allows us
to generate graphs with cluster structure, in which the different clusters exhibit potentially interesting
centrality distributions, which will serve as an interesting testbed to explore the clustering obtained from the
eigenvectors of our CGSOs.
36 R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

1.5
Degree 1.5
K-Core 12 1.5
Count of Walks 1.5
PageRank
7 6.8
1.0 1.0 1.0 6 1.0
10 6.6
6
5
0.5
5
0.5
8
0.5 0.5 6.4
4
4 6.2
e2

e2

e2

e2
0.0 0.0 0.0 0.0
6 3 6.0
0.5 3 0.5 0.5 0.5
4 2 5.8
2
1.0 1.0 1.0 1 1.0
5.6
1 2
1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 0 1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 5.4
e3 e3 e3 e3
Figure 3.1 – Result for the spectral clustering task on the Cora graph (Sen et al., 2008) with core numbers considered
as clusters. We report the values of the Adjusted Mutual Information (AMI) in percentage for different
combinations of the exponents (e2 , e3 ) in Ve2 AVe3 .

Experimental Setting. To better understand the information contained in the spectral decomposition
of our CGSOs we will now generate graphs from our SBBAMs and use the spectral clustering algorithm
defined on the basis of our CGSOs to attempt to cluster our generated graphs. In our experimental setting,
each block or BA graph has 100 nodes and an individual parameter r, specifically, r1 = 5, r2 = 10 and
r3 = 15. In addition we set pij = 0.1 for all i ̸= j with i, j ∈ {1, 2, 3}. Figure a.2 in Appendix a.5 gives an
example of an adjacency matrix sampled from this model. We observe variations in edge density across
different blocks and in particular observe homophilic cluster structure in the third block, while the first block
appears to be predominantly heterophilic, a rather challenging and interesting structure.
Figure a.3 in Appendix a.6 illustrates the k-core distribution of the three individual BA blocks and the
combined SBBAM. Notably, the k-core distribution distinguishes the three BA graphs, while the nodes in
the combined graph exhibit less discernibility by k-core.
Following the graph generation, we perform spectral clustering (see Algorithm 5 in Appendix a.7) using
our CGSOs to asses their ability to recover the blocks in our generated SBBAM. Specifically, we utilize
the three eigenvectors of Φ = Ve2 AVe3 corresponding to the three largest eigenvalues of different CGSOs
defined in Section 3.3.1. Working with this particular parametrized form our CGSOs further allows us to
study the effect of different centrality normalizations with e2 , e3 ∈ [−1.5, 1.5]. We repeated each experiment
200 times, and then reported the mean and standard deviation of Adjusted Mutual Information (AMI) and
Adjusted Rand Index (ARI) values. For consistency, we used the same 200 generated graphs for all the
GSOs and the baselines.
In Figure 3.1, we report the AMI values using the four centralities. As noticed, while having competitive
results between the degree centrality, the PageRank score and the count of walks, we reach the highest
AMI values by using the k-core centrality metrics. Using the degree centrality, we reach the highest AMI
value when both exponent e2 and e3 are negative, while for the k-core and the number of walks, we notice a
different behavior as the AMI increase when both the exponents e2 and e3 are positive. Thus, we conclude
that nodes with higher k-core and count of walks are important for this setup, i.e., when the node labels
are positively correlated with global centrality metrics such as the k-core. We report the ARI values of the
same experiment in Appendix a.8.

3.4.2 Centrality Recovery in Spectral Clustering

In this experiment, we aim to discern the CGSOs’ effectiveness in recovering clusters based on centrality
within a real-world graph. Using the Cora dataset, we chose core numbers to indicate centrality-based
clusters. We aim to assess the capacity of various CGSOs to effectively recover clusters reflective of
3.5 E X P E R I M E N TA L E VA L U AT I O N 37

Table 3.1 – The result of the spectral clustering task on the synthetic graph data. We present the mean and standard
values of AMI and ARI in percentage. ⃝ 1 Spectral clustering using the centrality based GSOs, ⃝
2 Other
baselines.

Method AMI in % ARI in %


Fast Greedy 17.27±4.28 19.98±5.03
Louvain 14.37±3.34 14.82±3.92

2
Node2Vec 1.11±0.92 1.17±0.96
Walktrap 1.39±1.16 1.14±0.97

CGSO w/ D 23.26±3.36 22.95±3.86


CGSO w/ Vcore 35.78±4.67 33.76±5.83

1
CGSO w/ Vℓ-walks 23.85±3.64 25.00±4.18
CGSO w/ VPR 35.62±4.90 33.19±6.03

core numbers. This investigation aims to shed light on their potential utility in capturing centralities and
hierarchical structures within intricate graphs.
Spectral Clustering on Cora. In this experiment, we consider only the largest connected component of
the Cora graph. We use the spectral clustering algorithm on the different CGSOs to recover K clusters,
where K is the number of possible core numbers in the graph. We repeat each experiment 10 times,
and report the average AMI and ARI values. We also compared our CGSOs with the popular Louvain
community detection method (Blondel et al., 2008), the node2vec node embedding methods (Grover and
Leskovec, 2016) combined with the k-means algorithm, the Walktrap algorithm (Pons and Latapy, 2005),
and the Fast Greedy Algorithm which also optimizes modularity by greedily adding nodes to communities
(Clauset, Newman, and Moore, 2004). For the walk count node centrality matrix, we used ℓ = 2 in all our
experiments. We consider the CGSO Φ = Ve2 AVe3 , where we normalize the adjacency matrix with the
topological diagonal matrix V using different exponents (e2 , e3 ).
The results of the spectral clustering on this synthetic graph are presented in Table 3.1. As expected,
normalizing the adjacency matrix with k-core yields higher AMI and ARI values. This observation indi-
cates an improved discernment of each node’s membership in its respective cluster, achieved through
the incorporation of global centrality metrics. Our CGSO outperforms well-known community detection
techniques, such as the Louvain algorithm, which optimizes the modularity, measuring the density of links
inside communities compared to links between communities. However, in our setting, some blocks have
fewer inter-edges than intra-edges with other blocks, thus making it difficult for the Louvain algorithm to
cluster these nodes using the edge density. This experiment further reinforces the intuition that if different
clusters exhibit different centrality distributions then our CGSOs are able to capture this difference better
than other clustering alternatives which leads to better clustering performance.

3.5 E X P E R I M E N TA L E VA L U AT I O N

We begin by discussing our experimental setup. Further details on the datasets we evaluate on and the
training set-up can be found in Appendix a.1.
Baselines. We experiment with two particular instances of our proposed CGNN model, using a GCN
and GATv2 as the backbone models, we refer to this instance as CGCN and CGATv2, respectively. We
38 R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

Table 3.2 – Classification accuracy (± standard deviation) of the models on different benchmark node classification
datasets. The higher the accuracy (in %) the better the model. ⃝1 GCN-based models ⃝ 2 Other vanilla
GNN baselines ⃝ 3 CGCN ⃝ 4 CGATv2. Highlighted are the first, second best results. OOM means Out of
memory.
Model CiteSeer PubMed arxiv-year chamelon Cornell deezer-europe squirrel Wisconsin

GCN w/ A 64.95±0.58 77.12±0.61 38.55±0.71 61.03±1.31 57.03±3.91 57.65±0.84 22.38±6.06 54.51±1.47


GCN w/ L 28.11±0.54 43.65±0.71 32.81±0.29 56.97±0.75 54.32±0.81 53.92±0.59 36.20±0.84 60.00±2.00
GCN w/ Q 63.28±0.80 76.57±0.59 33.76±2.36 53.88±2.35 35.41±2.55 56.79±1.79 27.69±2.21 53.33±0.78

1 GCN w/ Lrw 30.18±0.74 59.68±1.03 36.36±0.24 48.77±0.54 61.62±1.08 54.04±0.44 34.27±0.35 65.10±0.78
GCN w/ Lsym 29.90±0.66 57.68±0.45 36.49±0.14 50.81±0.24 60.27±1.24 53.30±0.45 35.96±0.28 66.08±2.16
GCN w/ Â 68.74±0.82 78.45±0.22 42.23±0.25 58.44±0.26 56.22±1.62 60.68±0.45 37.73±0.33 57.45±0.90
GCN w/ H 66.15±0.55 76.45±0.48 41.27±0.21 56.51±0.47 54.86±1.24 59.45±0.50 38.23±0.47 54.31±0.90

GIN 66.62±0.44 78.22±0.52 38.27±3.43 61.60±1.05 45.95±3.42 OOM 25.78±5.12 58.82±1.75


GAT 59.84±3.14 71.55±4.69 41.26±0.30 63.60±1.70 49.46±8.11 57.67±0.74 40.37±2.89 55.88±2.81

2 GATv2 63.01±2.97 73.96±2.22 41.16±0.25 64.14±1.53 43.78±4.80 56.77±1.19 42.63±2.61 53.53±4.12
PNA 48.89±11.15 70.83±6.51 32.45±2.34 22.89±1.09 40.54±0.00 OOM OOM 53.14±2.55

CGCN w/ D 68.35±0.45 78.70±1.10 45.39±0.45 64.17±8.10 72.43±13.09 58.04±1.06 42.30±1.34 76.86±7.70


CGCN w/ Vcore 68.40±0.75 77.91±0.41 47.27±0.31 63.68±5.00 73.78±12.16 60.90±2.28 40.59±2.21 74.90±6.52

3
CGCN w/ Vℓ-walks 67.31±0.75 77.57±0.37 39.35±0.49 66.21±2.49 72.70±3.24 59.15±1.24 36.03±5.81 74.90±4.19
CGCN w/ VPR 67.11±0.56 78.17±4.27 47.14±0.31 60.94±7.00 76.22±16.3 63.41±0.77 32.17±3.94 80.78±11.7

CGATv2 w/ D 68.60±0.60 77.46±0.51 45.09±0.17 58.22±2.74 76.49±4.37 OOM 35.30±2.32 85.69±3.17


CGATv2 w/ Vcore 68.83±0.66 77.99±0.43 44.38±0.25 55.83±2.28 75.95±3.72 OOM 34.17±1.45 85.10±2.80

4
CGATv2 w/ Vℓ-walks 68.11±0.91 75.43±0.89 46.70±0.21 55.59±2.57 74.32±5.70 OOM 34.25±2.15 83.53±2.66
CGATv2 w/ VPR 68.97±0.65 78.46±0.23 41.64±0.18 58.82±1.68 74.05±4.55 OOM 38.41±1.66 80.78±2.45

compared the proposed CGCN to GCN with classical GSOs: the adjacency matrix A, Unormalised
Laplacian L = D − A, Singless Laplacian Q = D + A (Cvetković and Simić, 2010), Random-walk
Normalised Laplacian Lrw = I − D−1 A, Symmetric Normalised Laplacian Lsym = I − D−1/2 AD−1/2 ,
Normalised Adjacency  = D−1/2 AD−1/2 (Kipf and Welling, 2017b) and Mean Aggregation H = D−1 A
(Xu et al., 2019b). We also compare to other standard GNN baselines: Graph Attention Network (GAT)
(Veličković et al., 2018), Graph Attention Network v2 (GATv2) (Brody, Alon, and Yahav, 2022), Graph
Isomorphism Network (GIN) (Xu et al., 2019b), and Principal Neighbourhood Aggregation (PNA) (Corso
et al., 2020).

3.5.1 Experimental Results

We present the performance of our CGCN and CGATv2 in Table 3.2. The performance of CSGC, i.e.
centrality based Simple Graph Convolutional Networks (Wu et al., 2019a), in Appendix a.3. We also
incorporated our learnable CGSOs into H2GCN (Zhu et al., 2020) resulting CH2GCN, that go beyond the
message passing scheme and which is designed for heterophilic graphs, we detailed the experiment and
the results in Appendix a.9. The results of CGCN, CGATv2, CSGC and the other baselines on additional
datasets can be found in Table a.4 of Appendix a.2, and Tables a.5 and a.6 of Appendix a.3. It has been
observed that, across numerous datasets, CGCN and CGATv2 outperform classical GSOs and vanilla
GNNs. Moreover, it is noteworthy that the optimal choice of centrality for CGCN varies depending on the
specific dataset. To better understand the choice of each centrality, we displayed the learned weights of
3.6 C O N C L U S I O N 39

Table 3.3 – Classification accuracy (± standard deviation) of the models on different benchmark node classification
datasets. The higher the accuracy (in %) the better the model.
Model CiteSeer PubMed arxiv-year chamelon Cornell deezer-europe squirrel Wisconsin
CGCN w/ D 68.35±0.45 78.70±1.10 45.39±0.45 64.17±8.10 72.43±13.09 58.04±1.06 42.30±1.34 76.86±7.70
CGCN w/ Vcore 68.40±0.75 77.91±0.41 47.27±0.31 63.68±5.00 73.78±12.16 60.90±2.28 40.59±2.21 74.90±6.52
CGCN w/ Vℓ-walks 67.31±0.75 77.57±0.37 39.35±0.49 66.21±2.49 72.70±3.24 59.15±1.24 36.03±5.81 74.90±4.19
CGCN w/ VPR 67.11±0.56 78.17±4.27 47.14±0.31 60.94±7.00 76.22±16.3 63.41±0.77 32.17±3.94 80.78±11.7

CGCN w/ D − Vcore 69.0±0.64 78.77±0.34 48.37±0.15 65.04±4.37 73.24±6.56 59.81±0.51 40.74±4.77 74.51±3.62
CGCN w/ D − Vℓ-walks 67.99±0.55 78.53±0.39 49.12±0.41 58.09±3.78 74.32±2.77 59.30±0.70 34.49±2.66 81.37±3.64
CGCN w/ D − VPR 68.45±0.6 77.75±0.55 39.63±1.27 64.32±3.13 72.97±4.98 59.28±0.75 42.80±6.58 74.31±3.97

CGCN together with some statistics of each dataset in Tables a.10, a.11, a.12 and a.13. Several trends
are clear: i) For all the centrality metrics, the exponent e1 is usually positive for most of the datasets,
which indicates that an additive normalization of the GSO with our centralities in-style of the unnormalized
Laplacian often leads to optimal graph representation. However, the exponent values e2 and e3 have different
behaviors across centrality metrics, e.g., when using the PageRank centrality, the exponents e2 and e3 are
almost null for the graph datasets that are strongly homophilous indicating that an unnormalized sum over
neighborhoods is optimal. ii) When using the PageRank and Count of walks centrality metrics, we notice
that the parameter a is always negative for non-homophilous datasets. This is a very interesting finding
indicating that a representation with negatively weighted self-loops is advantageous for non-homophilous
datasets (an observation that we have not previously seen in the literature). iii) For the datasets where the
k-core centrality performs well (i.e. Cornell, arxiv-year, Penn94, and deezer-europe), we notice that the
parameter m3 is very close to zero, i.e, the regularization by adding an identity matrix to the CGSO turns
out to be best-ignored in these settings. These findings suggest that the optimal GSO components vary
depending on the graph type, highlighting the need for adaptable CGSO approaches rather than relying
solely on classical GSOs.
General intuition on the choice of centrality that we can provide relates to the fact that the node degree is
a local centrality metric, while the remaining three centralities we consider correspond to global metrics.
Therefore, it is apparent that if the learning task only requires local information a degree-based normalization
of the GSO is likely beneficial, while global centrality metrics are appropriate if more global information
is required. Beyond this statement it seems to be difficult to provide general guidance on the choice of
the global centrality metrics. Therefore, including both local and global centrality-based CGSO in the
CGNN might be optimal to dynamically distinguish the best type of centrality. We present the results of this
experiment in Tables 3.3 and a.7. By combining local and global centralities in the CGNNs, we usually
increase their performance.

3.6 CONCLUSION

In this work, we have proposed CGSOs, a novel class of Graph Shift Operators (GSOs) that can leverage
different centrality metrics, such as node degree, PageRank score, core number, and the count of walks
of a fixed length. Furthermore, we have modified the message-passing steps of Graph Neural Networks
(GNNs) to integrate these CGSOs, giving rise to a novel model class the CGNNs. Experimental results
comparing our CGNN models to existing vanilla GNNs show the superior performance of CGNN on many
real-world datasets. These experiments furthermore allowed us to analyse the optimal parameters of our
CGSO, which led to new and interesting insight such as for example an apparent benefit of negatively
40 R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

weighted self-loops for non-homophilous graphs. To further understand the cases where each centrality is
beneficial, we conducted additional experiments focused on spectral clustering using two distinct types
of synthetic graphs. Through these experiments, we identified instances where CGSOs outperformed
conventional GSOs.
A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N
4
Raph Neural Networks (GNNs) have proven to be highly effective in various graph learning tasks. A

G key characteristic of GNNs is their use of a fixed number of message-passing steps for all nodes
in the graph, regardless of each node’s diverse computational needs and characteristics. Through
empirical real-world data analysis, we demonstrate that the optimal number of message-passing layers
varies for nodes with different characteristics. This finding is further supported by experiments conducted
on synthetic datasets. To address this, we propose Adaptive Depth Message Passing GNN (ADMP-GNN),
a novel framework that dynamically adjusts the number of message passing layers for each node, resulting
in improved performance. This approach applies to any model that follows the message passing scheme.
We evaluate ADMP-GNN on the node classification task and observe performance improvements over
baseline GNN models.

4.1 INTRODUCTION

A plethora of structured data comes in the form of graphs (Bornholdt and Schuster, 2001; Cao et al.,
2020). This has driven the need to develop neural network models, known as Graph Neural Networks
(GNNs), that can effectively process and analyze graph-structured data. GNNs have garnered significant
attention for their ability to learn complex node and graph representations, achieving remarkable success
in several practical applications (Castro-Correa et al., 2024a; Corso et al., 2022; Duval et al., 2023b;
Rampášek et al., 2022b). Many of these models are instances of Message Passing Neural Networks
(MPNNs) (Kipf and Welling, 2017a; Xu et al., 2019c). A common characteristic of GNNs is that they typically
employ a fixed number of message passing steps for all nodes, determined by the number of layers in the
GNN. This static framework raises an intriguing question: Should the number of message passing steps be
adapted individually for each node to better capture their unique characteristics and computational needs?
Determining the optimal number of message passing layers for each node in a GNN presents a significant
challenge due to the intricate and diverse nature of graph structures, node features, and learning tasks.
While deeper GNNs can capture long-range dependencies (Liu et al., 2021a), they can also encounter
issues like oversmoothing, where nodes become indistinguishably similar (Giraldo et al., 2023; Luan et al.,
2022). In dense graphs, where information can propagate quickly, even shallow GNNs can effectively
capture local information (Zeng et al., 2020). Conversely, sparse graphs, particularly those with isolated
nodes or limited connectivity, may require additional layers to facilitate effective information sharing (Zhang
et al., 2021a; Zhao and Akoglu, 2020). This underscores the importance of selecting the appropriate
number of layers for a GNN to capture the necessary graph information effectively. An even more compelling
idea is to adjust the GNN depth for each node based on its local structural properties. This adaptive
approach could be especially beneficial for graphs with varied local structures, ensuring that each node is
processed according to its unique requirements.
This need for per-node customization naturally aligns with the concept of Dynamic Neural Networks, also
referred to as Adaptive Neural Networks (Bolukbasi et al., 2017). Dynamic Neural Networks represent
a class of models that can adjust their architecture or parameters depending on the input. Dynamic
Neural Networks have gained significant popularity, especially in the field of computer vision. Examples of
adaptation include varying the number of layers and implementing skip connections (Huang et al., 2016; Li
et al., 2017; Sabour, Frosst, and Hinton, 2017). Beyond computer vision, other types of adaptive neural

41
42 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

networks have been explored in various domains. In natural language processing, adaptive computation
time models allow recurrent neural networks (RNNs) to determine the required number of recurrent steps
dynamically based on the complexity of the input sequence (Graves, 2016). However, applying these
adaptations to graph learning tasks presents unique challenges. Unlike the structured and homogeneous
data often encountered in computer vision, graph data involves overcoming nuisances related to the
inherent complex structure of graphs. Graphs may have varying node degrees, non-uniform connectivity,
and feature Heterogeneity. While some dynamic approaches may extend effectively to graph classification
tasks, applying these methods to node classification presents additional complexities. In node classification,
each input sample (node) is part of a larger interconnected system where information propagates through
edges, i.e., the prediction for one node often depends on the features of the other nodes and the structure
of the graph. As a result, specific techniques are needed to account for the dependencies and relational
information encoded in the graph structure.
In this work, we focus on the task of node classification by proposing ADMP-GNN, a novel approach
that dynamically adapts the number of layers for each node within a GNN. Our main contributions are as
follows:
— Node-specific depth analysis in GNNs. We demonstrate through empirical analysis that different
nodes within the same graph may require varying numbers of message passing steps to accurately
predict their labels. This finding underscores the importance of node-specific depth in GNNs.
— Adaptive message passing layer integration. We present ADMP-GNN, a novel approach that
enables any GNN to make predictions for each node at every layer. Training the GNN to predict
labels across all layers is a multi-task setting, which often suffers from gradient conflicts, leading to
suboptimal performance. To address this, we propose a sequential training methodology where layers
are progressively trained, and their gradients are subsequently frozen, thereby mitigating conflicts
and improving overall performance.
— Adaptive layer policy learning for node classification. We introduce a heuristic method to learn a
layer selection policy using a set of validation nodes. This policy is then applied to select the optimal
layer for predicting the labels of test nodes, ensuring that each node exits the GNN at the most
appropriate layer for its specific classification task.
— Model-agnostic flexibility. Our approach is model-agnostic and can be integrated with any GNN
architecture that employs a message passing scheme. This flexibility enhances the GNN’s perfor-
mance on node classification tasks, providing a significant improvement over traditional fixed-layer
approaches.

4.2 R E L AT E D W O R K

In this section, we review key developments in GNNs and dynamical neural networks, which form the
foundation of our work.

DY N A M I C A L N E U R A L N E T W O R K S . Dynamic neural networks are gaining significance in the field of


deep learning. Unlike static models with fixed computational graphs and parameters during inference,
dynamic networks adapt their structures or parameters based on varying inputs. This dynamic flexibility
gives models significant benefits, such as improved accuracy, enhanced computational efficiency, and
superior adaptability (Wang et al., 2020b; Zhou et al., 2020). A popular type of dynamic neural networks
includes those that dynamically adjust network depth based on each input. For instance, in natural language
processing, some adaptive large language models employ adaptive depth to optimize both inference speed
and computational memory usage of the Transformer architecture (Elbayad et al., 2020; Schuster et al.,
4.2 R E L AT E D W O R K 43

Computers
0.425
0.400
Accuracy 0.375
0.350 Dense Subgraph
0.325 Sparse Subgraph
0 2 4 6 8 10
Layers
Photo
0.35
Accuracy

0.30
0.25
0.20
0 2 4 6 8 10
Layers
Figure 4.1 – Effect of GCN’s depth on sparse and dense subgraphs. The figure shows the performance of GCNs
when varying layer depths, and comparing its effectiveness on both sparse and dense subgraphs.

2022; Vaswani et al., 2017). In the field of computer vision, there are studies that dynamically generate
filters conditioned on each input, enhancing flexibility without significantly increasing the number of model
parameters (Jia et al., 2016).
In Graph Machine Learning, dynamic adaptations have been proposed for GNN message passing.
These methods include dynamically determining which neighbors to consider at each layer, enabling more
flexible and adaptive message passing (Finkelshtein et al., 2024), or allowing nodes to react to individual
messages at varying times rather than processing aggregated neighborhood information synchronously
(Faber and Wattenhofer, 2024). Another line of work focuses on adapting normalization layers for each
node to enhance expressive power, generating representations that reflect local neighborhood structures
(Eliasof et al., 2024). Additionally, other related approaches use residual connections to mitigate issues
like oversmoothing (Chen et al., 2020b; Errica et al., 2023). To the best of our knowledge, in the field of
GNNs, there has been no prior work proposing adaptive depth for each node. However, several studies
have focused on combining all GNN layers. These works typically aim to adapt GNN architectures for
heterophilic graphs (Chien et al., 2020) and leverage information from higher-order neighbors (Xu et al.,
2018a). While combining GNN layers can be viewed as a form of depth-adaptive strategy, where the final
node representation is guided by the optimal intermediate hidden states, this approach remains static
because the same inference policy is applied uniformly across all nodes and learned layer aggregators
stay fixed after training.
44 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

4.3 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

In this section, we first present an empirical analysis highlighting the necessity for node-specific depth
in GNNs. Then, we introduce our ADMP-GNN. Our study includes experiments on both synthetic and
real-world datasets to illustrate the importance and potential benefits of this methodology.

4.3.1 Depth Analysis on Synthetic Graphs

This analysis aims to highlight the importance of employing a varying number of message passing steps
based on the specific characteristics of individual nodes. As outlined in the introduction, this approach is
particularly relevant for graphs where nodes exhibit diverse properties, such as varying local structures. In
this experiment, we investigate the effect of the number of message-passing layers on nodes with varied
levels of local neighborhood sparsity. To achieve this, we construct a graph by merging two subgraphs
extracted from real-world datasets, namely Computers and Photo (Shchur et al., 2018b). Both subgraphs
contain the same number of nodes, exhibit nearly identical homophily, and have equally distributed node
labels. The main difference between these subgraphs lies in their structure, as one subgraph is sparse while
the other is dense. Consequently, nodes within each subgraph share comparable structural characteristics.
Details on the construction of these synthetic datasets and visualizations of their adjacency matrices are
provided in Appendix b.2, including Fig. b.1.
We have trained L + 1 different GCN models, with a varying number of layers ℓ ∈ {0, . . . , L}, where
L represents the maximum depth, set to L = 10. Each GCN was trained on the entire synthetic graph,
composed of both sparse and dense subgraphs, but the performance was evaluated separately on the
individual subgraphs. This allows us to assess the impact of GNN depth on different types of local subgraph
structures. The results, presented in Fig. 4.1, reveal notable differences in behavior between the subgraph
types. As observed, in dense subgraphs, the test accuracy decreases at a faster rate, while in sparse
subgraphs, the drop in accuracy occurs later, typically around layers 2 or 3. Moreover, the optimal number
of layers differs between sparse and dense subgraphs. For instance, in the Computers dataset, the
highest accuracy is achieved at layer 2 for the sparse subgraph, while for the dense subgraph, the optimal
performance is reached at layer 0. Additionally, in the Photo dataset, we observe a distinct behavior
starting from layer 6, where the impact of GNN depth diverges between sparse and dense subgraphs. This
highlights the need to adapt the number of layers per node based on its characteristics.

4.3.2 Adaptive Message Passing Layer Integration

Building on the insights from the previous analysis, it is essential to develop a framework that can
efficiently adjust the GNN depth. For a maximum GNN depth L, a traditional approach would involve
training L + 1 different GNNs, each with a distinct number of layers ℓ, where ℓ ranges from 0 and L.
Subsequently, a policy must be established to determine the optimal GNN with the appropriate number of
layers for each individual node. However, training L + 1 GNNs separately can be computationally expensive.
A more efficient approach involves designing a single GNN with L + 1 layers that provide predictions at each
intermediate layer. To ensure equivalence to the previous approach, i.e., training L + 1 GNNs separately,
the computational graph for predictions at layer ℓ must match that of a standard GNN with ℓ layers, and the
classification performance at layer ℓ should yield results comparable to those of a conventional GNN with ℓ
layers.
To address these challenges, we introduce ADMP-GNN, an adaptation of a Message Passing Neural
Network with a maximum depth of L layers. The goal is to ensure that the computational graph and the
4.3 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N 45

Ex. Update Ex. Update Ex. Update Ex. Update

Agg. Ct. Update Agg. Ct. Update Agg.

Aggregate function Model weights trained at


Model weights trained at
Continuation Update function
Model weights trained at
Exit Update function Model weights trained at

Figure 4.2 – Illustration of ADMP-GNN, when the maximum GNN depth is L = 3.

performance of ADMP-GNN at a certain layer match that of traditional GNNs when trained and tested on
(ℓ)
the same number of layers. To achieve this, we incorporate an additional Update function, denoted as ϕEx
(ℓ)
(Ex stands for ‘Exit’), to directly predict node labels at a given layer ℓ, i.e., p(ℓ) . The function ϕEx is defined
as follows,    
(ℓ) (ℓ) (ℓ−1) (ℓ) e (ℓ) m(ℓ)
pv = ϕEx hv , mv = Softmax W v ,
(ℓ)
where W e (ℓ) ∈ Rd ×c is a learnable weight matrix, d(ℓ) is the dimension of the hidden representation at the
ℓ-th layer, and c is the number of classes. To obtain predictions at a deeper layer ℓ′ ≥ ℓ, we continue the
(ℓ)
message passing using another Update function ϕCt (Ct stands for ‘Continuation’),
 
(ℓ) (ℓ) (ℓ−1) (ℓ)
pv = ϕEx hv , mv ,
 
(ℓ+1) (ℓ) (ℓ−1) (ℓ)
hv = ϕCt hv , mv .

(0)
For ℓ = 0, we directly use the Exit Update function on the node features, i.e., mv = xv ,
   
(0) (0) (0) e (0) x v .
∀v ∈ V , pv = ϕEx mv = Softmax W

In Fig. 4.2, we illustrate the architecture of the proposed ADMP-GNN.

4.3.3 Training Scheme of ADMP-GNN

Our next objective is to train ADMP-GNN to predict node labels across all layers ℓ ∈ {0, . . . , L} simulta-
(ℓ)
neously. For each layer ℓ, let θℓ denote the weights of the function ψ(ℓ) ◦ ϕCt (·). Two distinct strategies for
training ADMP-GNN are explored:

1 . A G G R E G AT E L O S S M I N I M I Z AT I O N ( A L M ) . This straightforward approach minimizes the aggregate


loss over all layers. The total loss is formulated as,

LALM := arg min Ev∈V [S L (v)] (4.1)


θ
" #
L  
= arg min Ev∈V ∑ L pv (mv , θ0 , . . . , θℓ ), yv
(ℓ) (0)
(4.2)
θ0 ,...,θ L ℓ=0
46 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

(ℓ) (0) (ℓ) (ℓ)


where pv (mv , θ0 , . . . , θℓ ) = ϕEx (mv ) is the prediction for the node v at layer ℓ, and L is the Cross
Entropy Loss. This approach may encounter gradient conflicts, particularly for early layers involved in both
computation and back-propagation across upper layers.

2 . S E Q U E N T I A L T R A I N I N G ( S T ) . We have studied an alternative training setup where we progressively


train one GNN layer at a time, subsequently freezing each layer after training. More formally, the problem in
(4.1) can be tackled using dynamic programming as follows,
 
(ℓ) (0)
∀v ∈ V , Sℓ+1 (v) = L pv (mv , θ0⋆ , . . . , θℓ⋆ , θℓ+1 ), yv
+ Sℓ ( v )
θℓ⋆ = arg min Ev∈V [Sℓ (v)] ,
θℓ

(0)
where ∀v ∈ V , S0 (v) = L(ϕEx , yv ). For each intermediate layer ℓ < L, by training this layer on the node
(ℓ)
classification task, we obtain high-quality node representations {hv : v ∈ V }. These representations
are directly employed for predictions and serve as a robust foundation for the label predictions of the
subsequent layer ℓ + 1. Algorithm 3 offers a summary of the approach.
Comparison of ADMP-GNN Training Paradigms. To identify the optimal multi-task training configuration,
we evaluate how much of a performance drop we lose at each layer compared to the single-task setting
in GNNs. The comparative analysis of the three strategies for both GCN and GIN is detailed in Tables
4.1, b.4, b.7, and b.8. Our findings indicate that ADMP-GNN ST outperforms ADMP-GNN ALM. Notably,
the performance of ADMP-GNN ST is comparable to, or even exceeds, that of GNN when trained under
the single-task setting. Furthermore, ADMP-GNN ST exhibits a smaller standard deviation, suggesting
more consistent performance. These results underscore the effectiveness of the ST setting in addressing
gradient conflicts inherent in aggregate loss minimization (ALM). Furthermore, ADMP-GNN ST effectively
mimics the results of training L + 1 separate GNNs, each with a different number of layers, while requiring
only a single unified model. The next step, outlined in Section 4.3.5, is to learn a policy that selects the
optimal prediction layer for each node, completing the framework of ADMP-GNN.

T I M E C O M P L E X I T Y.
The training setup ST, where we sequentially train the deep ADMP-GNN, incurs
relatively higher time costs due to the need for L + 1 training iterations. However, in each iteration,
backpropagation is performed on a limited number of parameters, approximately equivalent to those in a
single message passing layer. Consequently, only a small number of epochs are required for each training
iteration. We report the training time of each approach in Table b.2 in Appendix b.3.

4.3.4 Empirical Insights into Node Specific Depth

We define Oracle Accuracy as the maximum achievable test accuracy under an optimal policy for
selecting layers. This is formally expressed as,

1  

(ℓ)
Aoracle = 1 yv ∈ {ybv : ℓ = 0, . . . , L} ,
|Vtest | v∈V test

(ℓ)
where Vtest is the set of test nodes, {ybv : ℓ = 0, . . . , L} represents the predictions for node v at each layer,
and 1(·) is the indicator function. Specifically, Aoracle corresponds to the accuracy obtained if, for each test
node, we could perfectly choose the layer that provides the correct prediction. Assuming such an optimal
4.3 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N 47

Table 4.1 – Comparison of ADMP-GCN training paradigms ALM and ST. These paradigms are also compared to
the single-task training setting to evaluate which approach most closely mimics the classical GCN under
single-task training. The best results for each dataset are bolded.
# Layers Training Paradigm Model Cora CiteSeer CS PubMed Genius Ogbn-arxiv
Single-task GCN 56.38±0.04 57.18±0.12 88.04±0.49 72.50±0.09 80.82±1.00 48.88±0.06
0
ADMP-GCN (ALM) 56.96±0.20 58.44±0.21 87.06±1.06 72.11±0.18 80.03±0.37 36.50±0.12
Multi-task
ADMP-GCN (ST) 56.38±0.06 57.17±0.09 87.27±1.29 72.48±0.14 80.17±0.79 48.86±0.03
Single-task GCN 76.90±0.14 69.68±0.06 91.74±0.80 76.63±0.13 80.23±0.37 55.21±0.50
1
ADMP-GCN (ALM) 75.67±0.18 70.12±0.04 90.55±0.74 73.74±0.16 80.13±0.29 39.54±1.44
Multi-task
ADMP-GCN (ST) 76.90±0.00 69.70±0.00 90.89±0.81 76.60±0.00 79.93±0.00 55.15±0.00
Single-task GCN 81.06±0.50 71.05±0.48 91.67±0.94 79.46±0.31 79.88±0.51 66.92±0.67
2
ADMP-GCN (ALM) 70.82±4.05 60.95±2.39 31.20±12.85 75.10±2.41 79.64±0.60 55.25±0.92
Multi-task
ADMP-GCN (ST) 80.73±0.33 71.33±0.40 91.49±0.66 79.02±0.21 80.06±0.11 66.51±0.65
Single-task GCN 79.14±1.58 66.33±1.35 89.80±0.87 78.50±0.68 80.00±0.04 67.33±0.55
3
ADMP-GCN (ALM) 68.64±5.14 51.30±6.74 47.82±12.98 74.17±1.98 80.04±0.10 56.22±0.50
Multi-task
ADMP-GCN (ST) 80.21±0.52 70.08±0.90 89.83±1.02 78.25±0.51 79.93±0.00 68.31±0.48
Single-task GCN 75.96±1.93 60.33±2.38 78.90±22.33 76.59±0.98 80.01±0.04 65.49±0.99
4
ADMP-GCN (ALM) 67.88±6.05 49.29±7.87 52.76±11.79 73.40±1.95 80.04±0.10 56.43±0.31
Multi-task
ADMP-GCN (ST) 81.05±0.49 67.99±0.74 88.86±0.70 75.27±1.02 79.93±0.00 69.29±0.74
Single-task GCN 70.09±4.01 57.40±3.43 77.96±15.24 74.32±3.66 80.01±0.04 63.04±1.33
5
ADMP-GCN (ALM) 67.68±7.04 50.14±7.65 54.53±10.61 73.27±1.51 80.04±0.10 56.66±0.31
Multi-task
ADMP-GCN (ST) 81.22±0.36 67.42±0.75 86.65±1.52 75.08±0.94 79.93±0.00 68.92±1.00

selection for each node, it represents the upper bound of accuracy achievable when employing adaptive
strategies to determine the best layer for each node.
Based on the findings presented in Table 4.2 (and Table b.5 in Appendix b.5) for GCN, as well as Table
b.9 (and Table b.10 in Appendix b.6), it is clear that Aoracle outperforms the highest accuracies achieved by
both GCN and ADMP-GCN ALM. This empirical evidence strongly indicates that adopting a distinct exit
layer for each node is highly beneficial in an optimal configuration.

Table 4.2 – Comparison of highest accuracy (± standard deviation) for GCN and ADMP-GCN ST, with the layer
achieving the best accuracy indicated in brackets. The final row shows the Oracle Accuracy for ADMP-
GCN ST. Best results for each dataset are bolded.
Model Cora CiteSeer CS PubMed Genius Ogbn-arxiv

GCN 81.06±0.50 [2] 71.05±0.48 [2] 91.67±0.94 [2] 79.46±0.31 [2] 80.82±1.00 [0] 67.33±0.55 [3]
ADMP-GCN (ST) 81.22±0.36 [5] 71.33±0.40 [2] 91.49±0.66 [2] 79.02±0.21 [2] 80.17±0.79 [0] 69.29±0.74 [4]
ADMP-GCN (ST) - Oracle 89.43±0.19 81.96±0.49 97.24±0.52 90.13±0.36 85.97±7.59 79.64±0.27
48 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

4.3.5 Generalization to Test Nodes

In this section, we introduce a heuristic approach to predict the optimal exit layer for test nodes based on
the assumption that nodes exhibiting structural similarity should share the same exit layer. The notion of
structural similarity can be assessed using various metrics. In our work, we define the structural similarity
based on node centrality metrics, i.e., nodes are considered structurally similar if they exhibit closely aligned
centrality values within the graph. We measured node centralities using both local metrics like degree and
global metrics such as k-core (Malliaros et al., 2020), PageRank scores (Brin and Page, 1998), and Walk
Count, which are detailed below.

k - C O R E . The k-core decomposition of a graph is a method that involves systematically pruning nodes.
In each iteration, a subgraph Gk ⊂ G is constructed by removing all vertices whose degree is less than
k. The core number of a vertex u, denoted as core(u), represents the largest value of k for which i is
included in the k-core. Formally, this is expressed as core(i ) = max{k : i ∈ Gk }. Nodes with higher core
numbers are typically found in denser subgraphs, reflecting strong interconnections with their neighbors
and indicating a higher level of centrality.

PA G E R A N K .A node’s PageRank score represents the probability of a random walk visiting that node,
making it a fundamental metric for measuring node importance in social network analysis and the web
(Brin and Page, 1998).

W A L K C O U N T. The Walk Count centrality quantifies the importance of a node based on the number of
walks of a given length ℓ starting from that node. This measure captures higher-order connectivity patterns
within the graph, going beyond direct neighbors to account for paths that traverse intermediate nodes. In
this work, we consider the Walk Count centrality for ℓ = 2, as it effectively captures higher-order structures
and provides valuable insights into the organization of complex systems (Benson, Gleich, and Leskovec,
2016).
Using this assumption, we employ node clustering to partition the node set into C clusters, denoted by
P = ∪1≤c≤C Pc . Each cluster c ∈ {1, . . . , C } is assigned a common exit layer ℓc ∈ {0, . . . , L}, determined
using nodes excluded from the test set. Validation nodes are utilized for this purpose. (i) Centrality scores
are computed for all nodes, considering metrics. (ii) Nodes are ranked based on their centrality scores. (iii)
These centrality scores are discretized into C equal-sized buckets to facilitate the clustering process. (iv)
The optimal exit layer for each cluster is determined by evaluating the classification accuracy on validation
nodes within that cluster, i.e.,
n o
(ℓ)
∀c ∈ {1, . . . , C }, ℓc = arg max Acc pv ∈ Vval ∩ Pc .
ℓ∈{0,...,L}

For the cluster-based layer selection policy, we considered alternative deep learning approaches to
predict the optimal exit layer for each node. This high accuracy leads to a significant distribution shift
between the predictions for train nodes and those for test nodes. Conversely, validation nodes, which were
not utilized during the training of the ADMP-GNN, present a more suitable option for learning the policy due
to their unbiased predictions. For the choice of the cluster-based layer selection policy, we could use a
variety of deep learning mechanisms to predict the best exit layer for each node. This included generalizing
the policy by training neural networks to identify layers that accurately predict node outcomes or by framing
the problem as an optimal stopping problem following the framework of (Huré et al., 2021). However, these
approaches require learning the policy on a set of nodes larger than the test set, which is impractical
4.4 E X P E R I M E N TA L R E S U LT S 49

Table 4.3 – Classification accuracy (± standard deviation) on different node classification datasets for the baselines
based on the GCN backbone. The higher the accuracy (in %), the better the model. Highlighted are the
first, second best results. OOM means Out of memory.
Model Cora CiteSeer CS PubMed Genuis Ogbn-arxiv

JKNET-CAT 79.52±1.16 [2] 69.69±0.05 [1] 91.23±1.26 [1] 77.63±0.59 [2] 81.46±0.10 [2] 68.54±0.57 [5]
JKNET-MAX 75.67±0.18 [1] 70.12±0.04 [1] 90.55±0.74 [1] 75.10±2.41 [2] 80.13±0.29 [1] 56.66±0.31 [5]
JKNET-LSTM 78.95±0.62 [0] 65.83±1.27 [0] 90.17±1.41 [2] 77.73±0.67 [0] OOM OOM
Residuals - GCNII 76.84±0.20 [1] 69.72±0.06 [1] 90.84±1.35 [1] 77.82±0.44 [2] 81.36±1.13 [2] 61.48±3.10 [4]
AdaGCN 75.08±0.27 [1] 69.58±0.19 [1] 89.62±0.51 [0] 76.40±0.12 [5] 79.85±0.00 [0] 22.06±1.67 [5]
GPR-GNN 79.91±0.43 [2] 69.21±0.81 [2] 91.42±1.12 [2] 79.00±0.39 [2] 81.04±0.41 [2] 68.03±0.23 [3]

GCN 80.78±0.72 [2] 71.25±0.72 [2] 92.20±0.00 [1] 79.32±0.41 [2] 80.76±1.05 [0] 64.37±0.43 [2]
ADMP-GCN 81.22±0.36 [5] 71.33±0.40 [2] 91.49±0.66 [2] 79.02±0.21 [2] 80.17±0.79 [0] 69.29±0.74 [4]

ADMP-GCN w/ Degree 81.03±0.53 71.10±0.50 91.26±0.59 78.71±0.39 80.73±1.00 69.59±0.28


ADMP-GCN w/ k-core 81.19±0.40 71.27±0.53 91.29±0.68 78.73±0.41 80.73±1.00 69.55±0.33
ADMP-GCN w/ Walk Count 81.14±0.41 71.14±0.50 91.19±0.64 78.64±0.60 80.68±0.97 69.55±0.34
ADMP-GCN w/ PageRank 81.05±0.46 71.00±0.29 91.09±0.99 78.69±0.56 81.12±1.41 69.60±0.29

for node classification tasks where the training and validation node sets available for policy learning are
typically smaller than the test set.

4.4 E X P E R I M E N TA L R E S U LT S

Table 4.4 – Classification accuracy (± standard deviation) on different node classification datasets for the baselines
based on the GIN backbone. The higher the accuracy (in %), the better the model. Highlighted are the
first, second best results. OOM means Out of memory.
Model Cora CiteSeer CS PubMed Genuis Ogbn-arxiv

JKNET-CAT 77.94±0.67 [2] 64.82±0.04 [1] 89.26±1.18 [1] 75.89±2.50 [2] OOM 60.23±0.37 [1]
JKNET-MAX 77.47±0.81 [1] 64.96±2.08 [2] 87.40±2.11 [0] 76.21±1.73 [0] OOM 51.84±4.01 [0]
JKNET-LSTM 77.39±1.39 [2] 64.52±2.02 [1] 87.27±3.66 [3] 75.90±1.63 [0] OOM OOM
GPR-GIN 76.83±1.22 [2] 66.43±1.15 [2] 88.15±1.53 [5] 77.27±0.87 [2] 80.82±0.42 [3] 63.05±0.44 [2]

GIN 77.73±0.99 [2] 65.23±1.45 [2] 90.29±0.99 [1] 76.05±1.14 [2] 80.78±1.03 [0] 60.70±0.15 [1]
ADMP-GIN 78.07±0.68 [2] 65.41±1.91 [2] 90.82±1.15 [1] 76.46±1.04 [4] 80.47±0.91 [0] 60.85±0.01 [1]

ADMP-GIN w/ Degree 78.12±0.70 66.82±0.89 90.70±0.79 76.07±1.27 81.68±0.70 60.85±0.01


ADMP-GIN w/ k-core 78.10±0.64 66.36±1.03 90.85±0.93 75.93±1.13 81.48±0.64 60.85±0.01
ADMP-GIN w/ Walk Count 78.19±0.68 65.79±0.81 90.77±0.91 76.72±0.76 81.21±0.58 60.85±0.01
ADMP-GIN w/ PageRank 77.78±0.83 67.08±0.51 90.72±0.91 76.63±0.87 81.89±1.04 60.85±0.01
50 A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

4.4.1 Experimental Setup

D ATA S E T S . We use thirteen widely used datasets in the GNN literature. We particularly used the citation
networks Cora, CiteSeer, and PubMed (Sen et al., 2008), the co-authorship networks CS (Shchur et al.,
2018b), the citation network between Computer Science arXiv papers Ogbn-arxiv (Hu et al., 2020a), the
Amazon Computers and Amazon Photo networks (Shchur et al., 2018b), the non-homophilous dataset
genius (Lim and Benson, 2021), and the disassortative datasets Chameleon, Squirrel (Rozemberczki,
Allen, and Sarkar, 2021), and Cornell, Texas, Wisconsin from the WebKB dataset (Lim et al., 2021a). More
details and statistics about the used datasets can be found in Appendix b.1. For the Cora, CiteSeer, and
Pubmed datasets, we used the provided train/validation/test splits. For the remaining datasets, we followed
the framework in (Lim et al., 2021a; Rozemberczki, Allen, and Sarkar, 2021).

Algorithm 3: Sequential Training for ADMP-GNN


Input: Graph G = (V , E , X ), number of layers L, node classification loss function L
1 foreach t ∈ {0, . . . , L } do
2 if t = 0 then
(0) (0)
3 Set hv ← mv = xv for all v ∈ V .;
(0) (0) (0)
4 Compute predictions at layer ℓ = 0, i.e., ∀v ∈ V , pv = ϕEx (hv ).;
(0) (0)
5 Train the weights of ϕEx to minimize L( pv ).;
(0)
6 Freeze the gradients of ϕEx .
7 else
( t −1) ( t −1) ( t −1) ( t −1)
8 Use ϕCt to update node representations: ∀v ∈ V , h̃v = ϕCt ( hv ).;
(t) ( t −1)
9 Aggregate the information for neighbor nodes: ∀v ∈ V , mv = ψ(t) ({h̃u : u ∈ N (v)}).;
(t) ( t −1) ( t −1) (t)
10 Compute predictions at layer ℓ = t, i.e., ∀v ∈ V , pv = ϕCt (h̃v , mv ).;
( t −1) (t) (t)
11 Train ϕCt , ψ(t) , and ϕEx to minimize L( pv ).;
( t −1) (t)
12 Freeze the gradients of ϕCt , ψ(t) , and ϕEx .

BASELINES. We compare our approach with architectures that combine all the hidden representations
of nodes to form a final node representation used for prediction. For each baseline model, we vary the
number of layers from 0 to 5, and we report in Table 4.3 the performance of the best number of layers with
respect to the test set. (i) This includes Jumping knowledge, which combines the nodes representation
of all layers using an aggregation layer, e.g., MaxPooling (JKMaxPool), Concatenation (JK-Concat), or
LSTM-attention (JK-LSTM) (Xu et al., 2018b). (ii) Residuals-GCNII uses an initial residual connection and
an identity mapping for each layer. The initial residual connection ensures that the final representation
of each node retains at least a fraction of α from the input layer (Chen et al., 2020b). (iii) GPR-GCN
combines adaptive generalized PageRank (GPR) scheme with GNNs (Chien et al., 2020). (iv) Ada-GCN,
which proposes an RNN-like deep GNN architecture by incorporating AdaBoost to combine the layers
(Sun, Zhu, and Lin, 2019). In contrast to the baselines, which aggregate the layers and determine the
best performance by performing a grid search over the number of layers L, we fix the maximum number
of layers to L = 5 for ADMP-GNN. This comparison gives a strong advantage to the baselines, as their
results are optimized for each dataset, while ADMP-GNN’s performance is evaluated under a fixed and
consistent setup.

I M P L E M E N TAT I O N D E TA I L S .
We train all the models using the Adam optimizer (Kingma and Ba, 2015b)
and the same hyperparameters. The GNN hyperparameters in each dataset were performed using a Grid
4.5 C O N C L U S I O N 51

search on the classical GCN; we detail the values of these hyperparameters in Table b.3 of Appendix b.4.
To account for the impact of random initialization, each experiment was repeated 10 times, and the mean
and standard deviation of the results were reported. The experiments have been run on both an NVIDIA
A100 GPU.

4.4.2 Experimental Results

Through extensive experiments on multiple datasets, we can better understand the scenarios in which
ADMP-GNN proves to be effective. As observed in Tables 4.3, 4.4, b.6, and b.11, a comparison between
ADMP-GCN and ADMP-GIN against their respective baselines, GCN and GIN, demonstrates consistently
higher accuracy for most datasets. Regarding the centrality-based layer selection policy, it becomes clear
that this policy is particularly efficient when graphs exhibit a wide range of local density and centrality among
nodes. Most importantly, no impactful drop in accuracy was observed with ADMP-GNN. It is important
to note that the suitability of a given centrality metric can vary significantly across datasets. For instance,
while PageRank is widely used, its distribution often forms a peak near zero due to the normalization
constraint where the sum of PageRank scores across nodes equals 1, making it challenging to form
meaningful clusters. In contrast, other centrality metrics like k-core and Walk Count are more adaptable in
datasets where clear clustering patterns emerge. For example, in datasets like Cora, Ogbn, and Photo, the
k-core centrality effectively highlights clusters, enabling an adaptive depth policy. For datasets like Texas
and Wisconsin, Walk Count effectively captures the structural diversity necessary for cluster-based layer
selection policy. These observations emphasize the importance of selecting appropriate centrality metrics
tailored to the specific properties of a dataset to maximize the effectiveness of ADMP-GNN.

4.5 CONCLUSION

In this work, we propose ADMP-GNN, a novel adaption of message passing neural networks that enables
to make predictions for each node at every layer. Additionally, we propose a sequential training approach
designed to achieve performance comparable to training multiple GNNs independently in a single-task
setting. Our empirical analysis highlights the importance of node-specific depth in GNNs to effectively
capture the unique characteristics and computational requirements of each node. Determining the optimal
number of message-passing layers for each node is a challenging task, influenced by factors such as the
complexity and connectivity of graph structures and the variability introduced by node features. Through our
experiments, we identified that node centrality can serve as a useful indicator for determining the optimal
layer for each node. To this end, we heuristically learn a layer selection policy using a set of validation
nodes, which is then generalized to test nodes. Extensive experiments across multiple datasets, particularly
those characterized by a diversity in local structural properties, demonstrate that ADMP-GNN significantly
enhances the prediction accuracy of GNNs, offering an effective solution to address the challenges of layer
selection and node-specific learning.
Part III

ROBUSTNESS OF GRAPH NEURAL NETWORKS


A D V E R S A R I A L R O B U S T N E S S O F G N N S : T H E O R Y, B O U N D S , A N D D E F E N S E
5
Neural Networks (GNNs) have demonstrated state-of-the-art performance in various graph

G
RAPH
representation learning tasks. Recently, studies revealed their vulnerability to adversarial attacks.
In this work, we theoretically define the concept of expected robustness in the context of attributed
graphs and relate it to the classical definition of adversarial robustness in the graph representation
learning literature. Our definition allows us to derive an upper bound of the expected robustness of
Graph Convolutional Networks (GCNs) and Graph Isomorphism Networks subject to node feature attacks.
Building on these findings, we connect the expected robustness of GNNs to the orthonormality of their
weight matrices and consequently propose an attack-independent, more robust variant of the GCN, called
the Graph Convolutional Orthonormal Robust Networks (GCORNs). We further introduce a probabilistic
method to estimate the expected robustness, which allows us to evaluate the effectiveness of GCORN on
several real-world datasets. Experimental experiments showed that GCORN outperforms available defense
methods.

5.1 INTRODUCTION

Graph-structured data is prevalent in a wide range of domains, motivating therefore the development of
neural network models that can operate on graphs, known as Graph Neural Networks (GNNs). GNNs have
emerged as a powerful tool for learning node and graph representations. Many GNNs are instances of
Message Passing Neural Networks (MPNNs) (Gilmer et al., 2017) such as Graph Isomorphism Networks
(GIN)(Xu et al., 2019c) and Graph Convolutional Networks (GCN)(Kipf and Welling, 2017a). These
models have been successfully applied in real-world applications such as molecular design (Kearnes et al.,
2016). In parallel to their success, it has been shown, particularly in the field of computer vision, that
deep learning architectures can be susceptible to adversarial attacks (Goodfellow, Shlens, and Szegedy,
2015a). These attacks, which are based on injecting small perturbations into the input, lead to unreliable
predictions, limiting therefore the applicability of these models to real-world problems. Similar to other
deep learning architectures, GNNs are also vulnerable to adversarial attacks. Recent studies (Dai et al.,
2018; Günnemann, 2022; Zügner, Akbarnejad, and Günnemann, 2018) have shown that a GNN can be
attacked by applying small structural perturbations to the input graphs. These attacks pose a threat to the
reliability of GNNs, particularly in safety-critical applications such as finance and healthcare. Consequently,
different attacks have been proposed to explore the robustness of GNNs. Concurrently, several studies
focus on developing methods to mitigate possible perturbation effects and improve the robustness of
MPNNs. The proposed methods include augmenting training data with adversarial examples and retraining
the model (Feng et al., 2019), pre-processing methods such as edge pruning (Zhang and Zitnik, 2020),
and more recently robustness certificates (Schuchardt et al., 2021). While the majority of these proposed
defenses focus on structural perturbations, only limited advances have been made for feature-based
adversarial attacks on graphs. Moreover, despite the large amount of research on the robustness of these
methods through empirical exploration, there has been limited progress in understanding the theoretical
robustness of GNNs.
In this work, we conduct a theoretical examination of the robustness of MPNNs subject to adversarial
attacks. First, we formally define the concept of “Expected Adversarial Robustness” for structure and
feature-based perturbations in the context of graphs defined in a metric space. We furthermore establish a

55
56 A DV E R S A R I A L R O B U S T N E S S O F G N N S : T H E O R Y, B O U N D S , A N D D E F E N S E

formal relation between our introduced expected robustness and the conventional adversarial robustness
formulation. Further, by analyzing the input sensitivity of the iterative message-passing schemes, we derive
an upper bound on the robustness of these models for both structural and feature-based attacks. Motivated
by our theoretical results, we propose a refined learning scheme, called Graph Convolutional Orthonormal
Robust Network (GCORN), for the GCN, to improve its robustness to feature-based perturbations while
maintaining its expressive power. In addition to our theoretical findings, we empirically evaluate the
effectiveness of our GCORN in both node and graph classification on commonly used real-world benchmark
datasets and compare GCORN to existing methods to defend against feature-based adversarial attacks.
Most empirical evaluations of the effectiveness of defense methods consider worst-case scenarios, hence
not taking into account the variability of attacks and their likelihood of occurrence. To overcome this
limitation, we propose a novel probabilistic method for evaluating the expected robustness of GNNs, which
is based on our introduced robustness definition. The method is model-agnostic and can hence be applied
to any architecture to estimate the local robustness. More realistic and comprehensive evaluation of the
effectiveness of defense approaches can hence be conducted.
Our main contributions are: (i) We define and theoretically analyze the expected robustness of MPNNs,
producing an upper-bound on their expected robustness, (ii) a novel approach (GCORN) for improving the
expected robustness of GCNs to feature-based attacks while maintaining their performance in terms of
accuracy, (iii) a theoretically well-founded, probabilistic and model-agnostic evaluation method, and (iv) an
empirical evaluation of our GCORN on benchmark datasets, demonstrating its superior ability compared to
existing methods to defend against feature-based attacks.

5.2 R E L AT E D W O R K

While most studies on attacking machine learning models focus on images (Goodfellow, Shlens, and
Szegedy, 2015a), work on discrete spaces such as graphs has also emerged. Analogous to images,
most existing graph-based attack methods frame the task as a search/optimization problem. For instance,
the targeted attack Nettack (Zügner, Akbarnejad, and Günnemann, 2018) utilized a greedy optimization
scheme while Zügner and Günnemann (2019) formulate the problem as a bi-level optimization task and
leverage meta-gradients to solve it. Zhan and Pei (2021) expanded this work through a black-box gradient
algorithm overcoming several limitations. Furthermore, Dai et al. (2018) proposed to use Reinforcement
Learning to solve the search problem and generate adversarial attacks. Node injection attacks (Chen
et al., 2022; Ju et al., 2023; Tao et al., 2021; Zou et al., 2021) have also proven effective, with attackers
introducing malicious nodes instead of modifying existing nodes or edges, affecting therefore the model’s
performance. In parallel, the field of defending against adversarial attacks on GNNs is still relatively
under-explored compared to that of image-based models. The majority of methods are primarily focused
on heuristic strategies. Similar to computer vision, robust training (Zügner and Günnemann, 2019) and
noise injection (Ennadir et al., 2024) have been used to improve the robustness of GNNs. Additionally,
low-rank approximation with graph anomaly detection (Ma et al., 2021) has been used to defend against
adversarial attacks. The GNN-Jaccard method (Wu et al., 2019b) pre-processes the adjacency matrix
to detect potential manipulation of edges. In the same context, GNN-SVD (Entezari et al., 2020) uses a
low-rank approximation of the adjacency matrix to filter out noise. Other methods such as edge pruning
(Zhang and Zitnik, 2020) and transfer learning (Tang et al., 2020a) have also been proposed. Finally,
Wang et al. (2020a) proposed a low-pass adaptation of the message passing to enhance robustness
while providing theoretical guarantees. Although these defense strategies have had some success, for
the majority of them, their heuristic nature results in defenses against specific types of attacks without any
guarantees on the model’s underlying robustness. As a result, these defenses may be susceptible to being
circumvented by future new advanced attacks. Consequently, the investigation of robustness certificates
5.3 E X P E C T E D A DV E R S A R I A L R O B U S T N E S S 57

(Bojchevski and Günnemann, 2019b; Gosch et al., 2023; Zügner and Günnemann, 2019) has gained
attention by providing attack-independent guarantees on the stability of the model’s predictions such as
randomized smoothing (Bojchevski, Gasteiger, and Günnemann, 2020).
While the majority of the existing work on defense strategies for GNNs has focused on structural
perturbations, there is far less work on addressing feature-based attacks. This represents a significant
gap in the literature as feature-based attacks on GNNs can be very effective. Seddik et al. (2022) propose
to add a node feature kernel to the message passing to enhance the robustness of GCNs. Additionally,
RobustGCN (Zhu et al., 2019) proposes to use Gaussian distributions as the hidden representations,
enabling the absorption of the impact of both structural and feature-based attacks. Finally, Liu et al. (2021b)
edited the message passing module using adaptive residual connections and feature aggregation, which
have been shown experimentally to enhance the model’s robustness against abnormal node features.
Robustness certificates have also been proposed for node feature-based attacks (Bojchevski, Gasteiger,
and Günnemann, 2020; Scholten et al., 2022).

5.3 E X P E C T E D A DV E R S A R I A L R O B U S T N E S S

In this section we mathematically define the concept of the expected robustness for a graph-based
function, such as a GNN. Let us consider three metric spaces with defined norms over the graph space
(A, ∥·∥A ), the feature space (X, ∥·∥X ) and the label space (Y, ∥·∥Y ). Let D be the underlying probability
distribution defined on (A, X, Y). Given a graph-based function f : (A, X) → Y, and some input
G = (A, X ) ∈ (G, X) with its corresponding label y ∈ Y where f (A, X ) = y, the goal of an adversarial
attack is to produce a perturbed graph G e and its corresponding features X e which are slightly different
from the original input ( G, X ) such that the predicted class of ( G,
e Xe ) is different from the predicted class of
( G, X ). The adversarial task is contingent on defining a similarity measure between the input graph and
the adversarially generated graph. Note that for un-attributed graphs, the distance in 2.38 aligns with the
commonly used edit distance on graphs which is a measure of similarity between two graphs quantifying
the minimal number of edges that need to be edited to convert one graph into another while taking into
account graph isomorphism. Based on this distance, we can mathematically formulate the adversarial
task as finding a perturbed attributed graph (Ge, X e ) within a specified budget ϵ such that f (Ge, X e ) = ye ̸= y
with dα,β ([G , X], [Ge, X
e ]) < ϵ. Moreover, given that in practice the attacker does not have access to the
ground-truth labels, we define an adversarial graph attack to be valid when f (Ge, X e ) ̸= f (G , X). We can
now define the expected vulnerability of a graph function as its likelihood to suffer from such attacks in
the input’s neighborhood defined by ϵ. Upper bounding this vulnerability allows us to quantify the model’s
expected robustness, we start by formulating the expected vulnerability of a graph function f as

Advϵ [ f ] = P(G ,X)∼DG ,X [(Ge, X


α,β e ) ∈ Bα,β (G , X, ϵ) : dY ( f (Ge, X
e ), f (G , X)) > σ ], (5.1)

with Bα,β (G , X, ϵ) = {(Ge, X


e ) : dα,β ([G , X], [Ge, X
e ]) < ϵ} for any budget ϵ ≥ 0. Additionally, dY can be any
defined distance in the output space Y and σ > 0. Here, we focus on real-valued output mappings and
consider the distance metric dY ( f (Ge, X e ), f (G , X)) = ∥ f (Ge, X
e ) − f (G , X)∥Y . This formulation is applicable
to both graph and node classification tasks. In node classification, the parameter σ determines the minimal
number of nodes that need to be successfully attacked to consider the attack to be adversarially successful
at the graph-level. This allows for flexibility in different scenarios, where in some cases, even a single
node’s label flip may be considered a threat, while in others, a limited number of changes can be tolerated.
In graph classification, the parameter σ acts as a threshold to the continuous softmax output above which
an attack is considered effective. We can now introduce the concept of expected robustness of a function
defined on graphs, such as a GNN.
58 A DV E R S A R I A L R O B U S T N E S S O F G N N S : T H E O R Y, B O U N D S , A N D D E F E N S E

Definition 5.3.1 (Expected Adversarial Robustness). Let dα,β be a graph distance on the spaces (G , X)
and dY be a distance on the space Y . The graph function f : (G , X) → Y is ((dα,β , ϵ), (dY , γ))–robust if its
α,β
vulnerability as defined in 5.1 can be upper-bounded by γ, i. [Link]ϵ [ f ] ≤ γ.

In Appendix c.2 (Proposition c.2.1) we show how, via the equivalence of metrics, expected robustness in
a given metric implies expected robustness in several other metrics.
Relating Expected Adversarial Robustness to Worst-Case Adversarial Robustness. Our introduced
formulation represents a broader perspective of the classical “worst-case” adversarial robustness, where
the attacker aims to identify a single adversarial attack representing “worst-case” losses within a predefined
budget and neighborhood. Our Expected Adversarial Robustness focuses on understanding the overall
behavior of the underlying graph-based function within the specified input neighborhood. This approach
offers a more comprehensive assessment of the model’s robustness. Nevertheless, we note that our
formulation encompasses the adversarial robustness as a special case since by definition these examples
are included in our considered neighborhood. In fact, by adjusting the hyper-parameter σ, we can isolate
these worst-case examples. Lemma 5.3.2 directly relates Definition 5.3.1 to the classical “worst-case”
adversarial robustness.

Lemma 5.3.2. Let dα,β be a defined graph metric on the metric spaces G , X. Let f : (G , X) → Y be
a graph-based function, we have the following result: If f is ((dα,β , ϵ), (dY , γ))–robust, then f is also
((dα,β , ϵ), (dY , γ))–“worst-case” robust.

The proof of Lemma 5.3.2 is provided in Appendix c.1. As a result, our forthcoming theoretical analysis,
which considers the general case, is equally applicable to worst-case adversarial examples, which will
be also validated experimentally in Section 5.6. We finally note that the advantages and pitfalls of the
generalization of “worst-case” to average robustness have also been studied by Rice et al. (2021a).

5.4 T H E E X P E C T E D R O B U S T N E S S O F M E S S A G E PA S S I N G N E U R A L N E T W O R K S

We now use our Expected Adversarial Robustness Definition 5.3.1 to derive an upper bound on the
expected robustness of the GCN, based on which we introduce our more robust GCN adaptation.

5.4.1 On the Expected Robustness of Graph Convolutional Networks

Our work primarily focuses on the theoretical analysis of GCNs within the broader context of MPNNs. The
computations of one GCN layer are composed of the aggregation of node hidden states over neighborhoods
in the graph and subsequent node-wise updates of the hidden states via a weight matrix and non-linear
activation function.
The updated hidden states are then passed to the next layer for further aggregation and updates. An
iteration of this process can be expressed as follows

h(ℓ) = ϕ(ℓ) (Ah


e (ℓ−1) W (ℓ) ), (5.2)

where h(ℓ) represents the hidden state in the ℓ-th GCN layer and h(0) is the initial node features X ∈ Rn×K ,
W (ℓ) ∈ R p×e is the weight matrix in the ℓ-th layer, e is the embedding dimension and ϕ(ℓ) is a 1-Lipschitz
continuous non-linear activation function. Moreover, A e ∈ Rn×n , with n being the number the nodes, denotes
the normalized adjacency matrix A e =D − 1/2 AD − 1/2 .
Determining the exact expected adversarial robustness of a graph-based function, as outlined in Section
5.3, poses a significant challenge. To overcome this, we provide an upper bound, referred to as γ in
5.4 T H E E X P E C T E D R O B U S T N E S S O F M E S S A G E PA S S I N G N E U R A L N E T W O R K S 59

Definition 5.3.1. Our definition of adversarial attacks is closely related to the concept of sensitivity analysis.
In both cases, the goal is to understand how small input changes can affect the model’s output. We hence
tackled the adversarial task by adopting an input perturbation perspective. Similar approaches based on
sensitivity analysis have had some success for deep neural networks (DNNs). While extending to other
domains such as Convolutional Neural Networks is direct, generalizing to graphs presents new challenges.
Notably, model dynamics differ due to the Message Passing process involving the adjacency and node
features, complicating the task of providing an upper-bound. Since, the architecture itself involves the
adjacency matrix, any perturbation on the input produces a different dynamic in the model itself, unlike
the classical DNNs architecture, which remains static when subject to perturbations. Consequently, the
theoretical analysis and results have to reflect the underlying propagation architecture, i.e., the graph
structure. Theorem 5.4.1 provides theoretical insight into the expected robustness of GCNs by establishing
an upper bound on the amount of perturbation that a GCN can tolerate before its predictions become
unreliable.
Theorem 5.4.1. Let f : G → Y denote a graph-based function composed of L GCN layers, with W (i)
denoting the weight matrix of the i-th layer. Further, let d0,1 be a feature distance. For attacks targeting
node features of the input graph, with a budget ϵ, with respect to Definition 5.3.1:
L (ℓ)
— f is ((d0,1 , ϵ), (d1 , γ))–robust with γ = ∏ℓ= 1 ∥W ∥1 ϵ ( ∑u∈V ŵu ) /σ, with ŵu denoting the sum of
normalized walks of length ( L − 1) starting from node u and V is the node set.
L (ℓ)
— f is ((d0,1 , ϵ), (d∞ , γ))–robust with γ = ∏ℓ= 1 ∥W ∥∞ ϵ ŵ G /σ, with ŵ G = max ŵu .
u∈V

As previously mentioned, the provided upper-bound in Theorem 5.4.1 is directly dependent on both the
graph structure (in terms of walks from the graph’s nodes) and the propagation scheme (where the length
of the considered walks depends on the number of message-passing iterations). The derived upper-bound
is intuitive: effectively using feature-based attacks is increasingly difficult with increasing graph sparsity;
or conversely, in dense graphs node feature attacks can have greater effect since the message passing
scheme propagates them along a greater number of walks. To make this precise, the expected sum of
normalized walks in a sparser graph tends to be lower than in a denser one, resulting in a reduced bound,
indicating a more robust model within the considered neighborhood. We finally note that this bound applies
to both targeted and untargeted feature modifications, whether limited to a specific subset or all the nodes.
While our main focus is on node feature-based attacks, our analysis provided in Theorem 5.4.1 can be
extended to structural attacks; Theorem 5.4.2 sheds light on this latter case.
Theorem 5.4.2. Let f : G → Y denote a graph function composed of L GCN layers, where W (i) denotes
the weight matrix of the i-th layer. Further, let d1,0 be a graph distance. For structural attacks with a budget
ϵ, the function f is ((d1,0 , ϵ), (d2 , γ))–robust with
L L
γ= ∏ ∥W (ℓ) ∥2 ∥X∥2 ϵ(1 + L ∏ ∥W (ℓ) ∥2 )/σ. (5.3)
ℓ=1 ℓ=1

We observe the upper bound in 5.3 to functionally depend on the size of the graph via ∥ X ∥2 , yielding the
intuitive result that larger graphs have more potential targets to attack and thereby give rise to less robust
models. The proofs of Theorems 5.4.1 and 5.4.2 are provided in Appendix c.3.

5.4.2 On the Generalization to Other Graph Neural Networks

While our work focuses on GCNs, our analysis can be extended to any GNN. Our theoretical analysis is
contingent on assuming the input node feature space to be bounded, which is a realistic assumption for
real word data. For illustration, Theorem 5.4.3 derives the upper-bound of the specific case of GINs.
60 A DV E R S A R I A L R O B U S T N E S S O F G N N S : T H E O R Y, B O U N D S , A N D D E F E N S E

Theorem 5.4.3. Let f : (G , X ) → Y be composed of L GIN-layers (with parameter ζ = 0, that is usually


denoted by ϵ in the literature) and W (i) denote the weight matrix of the i-th MLP layer. We consider the
input node feature space to be bounded, i.e., ∥ X ∥2 < B for some B ∈ R, and graphs of maximum degree
∆G . For node feature-based attacks, with a budget ϵ, the function f is ((d0,1 , ϵ), (d∞ , γ))–robust with
L
γ= ∏∥W (l) ∥∞ ( B L ∆G + ϵ)/σ.
l =1

Theorem 5.4.3 is proved in Appendix c.4. In addition, following the same assumption, it is possible to
derive an upper bound on the GCN’s robustness when subject to both structural and feature-based attacks
simultaneously. The complete study and additional information are provided in Appendix c.6.

5.4.3 Enhancing the Robustness of Graph Convolutional Networks

Leveraging the established upper-bound on the expected robustness in Theorem 5.4.1, we now introduce
a novel approach, called Graph Convolutional Orthonormal Robust Networks (GCORNs), which enhances
the robustness of a GCN to node feature perturbations while maintaining its ability to learn accurate node
and graph representations. To this end, we aim to design a GCN architecture for which the upper bound γ
as stated in Theorem 5.4.1 is low. As this quantity is dependent on the norm of the weight matrices in each
layer of the GCN, we propose to control the norms of these matrices.
Enhancing the robustness through enforcing matrix norm constraints during training has previously been
studied in the context of DNNs through methods such as Parseval regularization (Cisse et al., 2017) or
optimizing over the orthogonal manifold (Ablin and Peyré, 2022). However, as observed in our experiments,
the added constraints of these methods can negatively impact the performance of the graph function
and additionally the hyper-parameter tuning can be tricky especially when dealing with large datasets. In
our work, we choose to tackle the task by modifying the mathematical formulation of the GCN in 5.2 to
encourage the orthonormality of the weights. We have therefore chosen to use an iterative algorithm (Björck
and Bowie, 1971), that mainly was used in the literature for studying Lipschitz approximation (Anil, Lucas,
and Grosse, 2019), to ensure a fair trade-off between clean and attacked accuracy. We note that, based on
the introduced Theorem 5.4.1, any orthonormalization method can theoretically enhance the underlying
model’s robustness. Given our weight matrix W, the iterative process which is computed using Taylor
expansion, consists of finding the closest orthonormal matrix Ŵ to our weight matrix W. By considering
Ŵ0 = W, we recursively compute Ŵk from
 
p
Ŵk+1 = Ŵk I + 12 Qk + . . . + (−1) p (−1/2
p ) Q k , (5.4)

with k ≥ 0, Qk = I − ŴkT Ŵk and p ≥ 1 is the chosen order. A key advantage of this orthonormalization
approach is its differentiability, hence its compatibility with the back-propagation process of training GNNs.
By incorporating this projection into each forward pass during the model’s training, we encourage the
orthonormality of the weights, consequently enhancing its expected robustness.
Training and Convergence of Our Approach. Encouraging the orthonormality of the weight matrices
for each layer in our framework mitigates the problems of vanishing and exploding gradients. Our proposed
method preserves the gradient norm, leading to enhanced convergence and improved learning (Guo et al.,
2022). It is important to note that the training of our framework relies on the convergence of the iterative
orthonormalization process. From the original work (Björck and Bowie, 1971) which examined this aspect,
convergence is contingent on the condition ∥W T W − I ∥≤ 1 being satisfied, which can be guaranteed by
applying a scaling factor, based on the spectral norm, to the weight matrices before the iterative process.
We noticed that this approach not only guarantees that the weight matrices satisfy the necessary condition
5.5 E S T I M AT I O N O F O U R R O B U S T N E S S M E A S U R E 61

for convergence but also helps to speed up the convergence and the training process as analyzed by
previous work (Salimans and Kingma, 2016).
Complexity of Our Approach. Equation 5.4 entails a trade-off between convergence speed and
the approximation’s precision. A higher order p yields closer projections in each iteration, resulting in
an increased computational complexity. This trade-off must be carefully considered to strike a balance
between complexity and approximation accuracy. The main complexity of the method results from the
matrix products which are O(e3 ) where e represents the embedding dimension. Our experiments indicate
that a low order and a small number of iterations often yield a satisfactory approximation in practice. We
note that our method’s complexity does not increase with growing input graph size, distinguishing it from
other defenses such as GNNGuard which has a complexity of O(e × | E|), where | E| represents the number
of edges or GNN-SVD which computes a low-rank approximation through SVD with a corresponding
complexity of O(n3 ), with n being the number of nodes. We point out that we conducted an ablation
study (reported in Appendix c.7) on the effect of changing the order p and number of iterations k on both
the precision of the orthonormal projection (represented by both the clean accuracy and the adversarial
accuracy) and the time complexity.

5.5 E S T I M AT I O N O F O U R R O B U S T N E S S M E A S U R E

Based on our introduced definition of expected robustness in Section 5.3, we introduce an algorithm to
α,β
empirically estimate the quantity Advϵ [ f ], serving therefore as a new metric to evaluate GNN robustness.
While most defense methods on graphs are evaluated using the worst-case accuracy of a GNN when
exposed to specific attack schemes (Bojchevski, Gasteiger, and Günnemann, 2020; Xu et al., 2019a;
α,β
Zügner and Günnemann, 2019), Advϵ [ f ] is an attack-independent and model-agnostic robustness metric
based on uniform sampling. It can therefore be used as a security checkpoint, in addition to existing
robustness metrics, to evaluate the GNN robustness against unknown attack distributions.
Numerous model-agnostic robustness metrics (Cheng et al., 2021; Weng et al., 2018a) have emerged
primarily in computer vision. These metrics mainly assess robustness through measuring the distortion
between input and output manifolds. While these methods can be extended to the graph classification
task with appropriate distortion definitions, their application in the node classification context proves
challenging. To our knowledge, there exists no specific or analogous adaptation within the context of graph
representation learning. Thus, our proposed evaluation aims to address this gap.
Given that the expected value of an indicator random variable for an event E is the probability of that
event, i.e., E[1{ E}] = P( E), we use 5.5 as an equivalent formulation of Advϵ [ f ] in 5.1.
α,β

h i
Advϵ [ f ] = E (G ,X)∼DG ,X ,
α,β
1{dY ( f (Ge, X
e ), f (G , X)) > σ } . (5.5)
(Ge,X
e )∈ Bα,β ((G ,X),ϵ)

From the law of large numbers (Bernoulli, 1713), the mean sampling is an unbiased estimator of the vulner-
α,β
ability quantity Advϵ [ f ]. Consequently, we estimate a GNN’s expected robustness by generating multiple
graph pairs ([G , X], [Ge, X
e ]), where each input graph [G , X] is sampled from the underlying distribution DG ,X
and [Ge, X
e ] is uniformly sampled from the ball of radius ϵ centered around [G , X], i. e. dα,β ([G , X], [Ge, X
e ]) ≤ ϵ.
This approach can be used for both structural and feature attacks. Since we focus on the latter, we propose
a sampling strategy based on the distance

d0,1 ([G , X], [Ge, X


e ]) = ∥X − X
e ∥X = max ∥Xi − X
ei∥p, (5.6)
i ∈{1,...,n}

where ∥·∥ p is the L p -norm (p > 0) and Xi , X e ∈ Rn × K .


e i are the i-th rows of the matrices X, X
62 A DV E R S A R I A L R O B U S T N E S S O F G N N S : T H E O R Y, B O U N D S , A N D D E F E N S E

Since the original dataset consists of samples [G , X] from the underlying distribution DG ,X , we only need to
sample uniformly graphs [Ge, X e ] from the neighborhood of the dataset graphs [ G, X ] such that ∥X − X e ∥X ≤ ϵ
e is equivalent to first sample Z ∈ Rn×K from Bϵ = Z ∈ Rn×K : ∥Z∥X ≤ ϵ

and Ge = G . Sampling such X
and then set Xe = X + Z. Sampling from the ball Bϵ could be performed by partitioning the ball with respect
to the possible values of ∥ Z ∥X .

Sr = { Z ∈ R n × K : ∥ Z ∥ X = r } , B ϵ = ∪ r ≤ ϵ Sr ; ∀r ̸ = r ′ Sr ∩ Sr′ = ∅.

We therefore propose to use the following Stratified Sampling approach, which is based on two main
steps: (i) Sample a distance r from [0, ϵ] using the distribution pϵ defined in Lemma 5.5.1; (ii) Sample Z
from Sr . Additional details on the prior distribution r and the sampling from Sr are in Appendix c.8.

Lemma 5.5.1. Let RK be the real finite-dimensional space and ϵ a positive real number. If R( p) is
the random variable indicating the maximum of the L p norm’s values inside the ball of radius ϵ, i.e.,
Bϵ = Z ∈ Rn×K : maxi∈{1,...,n} ∥Zi ∥ p ≤ ϵ . Then, for every p > 0, the density distribution of R( p) does not

 K −1
depends on p and is defined as follows, pϵ (r ) = K 1ϵ ϵr 1 {0 ≤ r ≤ ϵ }.

The proof is available in Appendix c.9. We note that the proposed framework is a practical and theory-
based robustness evaluation approach applicable to any GNN, since no assumption has been made on the
underlying architecture. Algorithm 6 in Appendix c.8 offers a summary of the approach.

5.6 E M P I R I C A L I N V E S T I G AT I O N

This section examines the practical impact of our theoretical findings on real-world datasets where we
aim to investigate the robustness of GCORN in comparison to other defense benchmarks.

5.6.1 Experimental Setup

The necessary code to reproduce all our experiments is available on github [Link]
In this section, we focus on node classification while we report results on graph classification (Morris et al.,
2020a) in Appendix c.10. We use the citations networks Cora, CiteSeer, and PubMed (Sen et al., 2008),
the Co-author network CS (Shchur et al., 2018b), and the citation network between Computer Science
arXiv papers OGBN-Arxiv (Hu et al., 2020b). Further information about the datasets and implementation
details are provided in Appendix c.12.
To reduce the impact of initialization, we repeated each experiment 10 times and used the train/vali-
dation/test splits provided with the datasets and for the CS dataset, we followed the framework of Yang,
Cohen, and Salakhudinov (2016). We note that for the node feature-based attacks, we normalized the
input features to work in a continuous space allowing more flexibility in terms of available attacks.
Attacks. We use two evasion and one poisoning feature-based attacks: (i) The baseline random attack
injecting Gaussian noise N (0, I) to the features with a scaling parameter ψ controlling the attack budget;
(ii) The white-box Proximal Gradient Descent (Xu et al., 2019a), which is a gradient-based approach to the
adversarial optimization task. We set the perturbation rate to 15%. While this attack has limited success for
structural perturbations due to the discrete space, it is known to be powerful in continuous spaces, such as
the node feature space;
(iii) The targeted-attack Nettack (Zügner, Akbarnejad, and Günnemann, 2018), where similar to the
original study, we selected 40 correctly classified target nodes, comprising 10 with the largest classification
margin, 20 randomly chosen, and 10 with the smallest margin.
5.6 E M P I R I C A L I N V E S T I G AT I O N 63

Table 5.1 – Attacked classification accuracy (± standard deviation) of the models on different benchmark node
classification dataset after the feature attacks.

Attack Dataset GCN GCN-k AirGNN RGCN ParsevalR GCORN

Cora 68.4 ± 1.9 69.2 ± 2.6 73.5 ± 1.9 71.6 ± 0.3 72.9 ± 0.9 77.1 ± 1.8
CiteSeer 57.8 ± 1.5 62.3 ± 1.2 64.6 ± 1.6 63.7 ± 0.6 65.1 ± 0.8 67.8 ± 1.4
Random
PubMed 68.3 ± 1.2 71.2 ± 1.1 70.9 ± 1.3 71.4 ± 0.5 71.8 ± 0.8 73.1 ± 1.1
(ψ = 0.5)
CS 85.3 ± 1.1 86.7 ± 1.1 87.5 ± 1.6 88.2 ± 0.9 87.6 ± 0.6 89.8 ± 1.2
OGBN-Arxiv 68.2 ± 1.5 52.8 ± 0.5 66.5 ± 1.3 63.8 ± 1.9 68.3 ± 1.9 69.1 ± 1.8

Cora 41.7 ± 2.1 46.3 ± 2.8 53.7 ± 2.2 52.8 ± 1.6 55.3 ± 1.2 57.6 ± 1.9
CiteSeer 38.2 ± 1.3 45.3 ± 1.4 49.8 ± 2.1 43.7 ± 2.2 51.2 ± 1.2 57.3 ± 1.7
Random
PubMed 60.1 ± 1.7 62.3 ± 1.3 62.4 ± 1.2 61.9 ± 1.2 61.3 ± 1.7 65.8 ± 1.4
(ψ = 1.0)
CS 69.9 ± 1.3 73.2 ± 0.9 76.7 ± 2.8 76.2 ± 1.4 78.7 ± 1.2 81.3 ± 1.6
OGBN-Arxiv 66.4 ± 1.9 46.6 ± 0.6 62.7 ± 1.6 63.0 ± 2.4 66.1 ± 0.7 67.3 ± 2.1

Cora 54.1 ± 2.4 58.3 ± 1.6 68.2 ± 1.8 62.5 ± 1.2 68.6 ± 1.7 71.1 ± 1.4
CiteSeer 52.3 ± 1.1 59.6 ± 1.6 59.3 ± 2.1 61.9 ± 1.1 62.1 ± 1.5 65.6 ± 1.4
PGD PubMed 66.1 ± 2.1 67.3 ± 1.3 70.8 ± 1.7 69.5 ± 0.9 68.9 ± 2.1 72.3 ± 1.3
CS 71.3 ± 1.1 74.1 ± 0.8 76.3 ± 2.1 76.6 ± 1.2 77.3 ± 0.6 79.6 ± 1.2
OGBN-Arxiv 67.5 ± 0.9 49.9 ± 0.7 55.7 ± 0.9 63.6 ± 0.7 67.6 ± 1.2 68.1 ± 1.1

Cora 60.9 ± 2.5 64.2 ± 5.2 66.7 ± 3.8 63.4 ± 3.8 67.5 ± 2.5 68.3 ± 1.4
CiteSeer 55.8 ± 1.4 71.7 ± 1.4 67.5 ± 2.5 70.8 ± 3.8 69.2 ± 3.8 77.5 ± 2.5
Nettack PubMed 60.0 ± 2.5 65.8 ± 2.9 69.2 ± 1.4 71.7 ± 3.8 68.3 ± 1.4 70.8 ± 1.4
CS 55.8 ± 1.4 71.6 ± 1.4 76.7 ± 1.4 71.7 ± 2.9 75.8 ± 2.8 78.3 ± 1.4
OGBN-Arxiv 49.2 ± 2.9 53.3 ± 1.4 56.7 ± 1.4 52.6 ± 2.5 55.8 ± 1.4 55.8 ± 1.4

Baseline Models. Currently, as presented in Section 5.2, there are only few methods for defending
against feature-based adversarial perturbations. We compare against (i) GCN-k (Seddik et al., 2022); (ii)
RobustGCN (RGCN) (Zhu et al., 2019); (iii) AIRGNN (Liu et al., 2021b) and (iv) we finally compared to the
Parseval Regularization (ParsevalR) (Cisse et al., 2017), another orthonormalization method, with great
success in the computer vision. The aim is mainly to motivate the merits of our proposed orthonormalization
method based on the iterative weight projection.

5.6.2 Experimental Results

Empirical Estimation of Our Expected Robustness. We start by investigating the robustness of the
α,β
different considered benchmarks through the empirical estimation of Advϵ [ f ] as introduced in Section 5.5
α,β
(see Appendix c.12). Figure 5.1 (a) and (b) report the estimated values of the vulnerability Advϵ [ f ] for our
GCORN and the considered defense benchmarks for the Cora and OGBN-Arxiv datasets, respectively.
Additional results on further datasets are provided in Appendix c.10. The figures show that our GCORN
64 A DV E R S A R I A L R O B U S T N E S S O F G N N S : T H E O R Y, B O U N D S , A N D D E F E N S E

1.0 1.0 0.8 GCORN

(a) (b) (c) GCN

0.8 0.8 0.6

Acc. S(ra, rd)


0.6 0.6
Adv[f]

Adv[f]
GCORN RGCN GCORN RGCN
GCN
GCN K
AIRGNN GCN
GCN K
AIRGNN 0.4
0.4 0.4
0.2 0.2 0.2

0.0 0.0 0.00


0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.0 0.1 0.2 0.3 0.4 0.5 0.6 5 10
ra (dotted), rd (solid)
α,β
Figure 5.1 – (a) and (b) display Advϵ [ f ] for Cora and OGBN-Arxiv. (c) Robustness guarantees on Cora, where
r a , rd are respectively the maximum number of adversarial additions and deletions.

yields the lowest adversarial vulnerability indicating therefore that our proposed approach is exhibiting
greater expected robustness within the considered neighborhood.
Worst-Case Adversarial Evaluation. Table 5.1 reports the attacked node classification accuracies
for our GCORN and the other considered benchmarks. The results show that, when subject to features-
based adversarial attacks with a varying budget, the performance of GCN is severely impacted. Instead,
the proposed GCORN demonstrates a significant improvement in defending against these attacks in
comparison to other methods. For instance, GCORN is outperforming the GCN-k by an average of ∼ 12%
in accuracy. In certain situations, GCORN demonstrates the ability to recover the GCN’s performance to
a level equivalent to no adversarial attack. Furthermore, the baselines are not as effective in defending
against gradient-based attacks. In contrast, GCORN, which is based on an attack-agnostic upper bound
(Theorem 5.4.1), demonstrates its ability to defend against gradient-based attacks. We, additionally study
the trade-off between robustness and input sensitivity in Appendix c.10.
Certified Robustness. In addition to our probabilistic and empirical evaluations, we assessed the
certified robustness of a GCN and our proposed GCORN on the Cora dataset using the sparse randomized
smoothing approach (Bojchevski, Gasteiger, and Günnemann, 2020). Specifically, we plotted their certified
accuracy S(r a , rd ) for varying addition r a and deletion rd radii. Figure 5.1(c) showcases our theoretical
robustness; GCORN improves the certified accuracy with respect to feature perturbations for all radii.
Structural Perturbations. Although our focus is on node feature-based adversarial attacks, we also
provided a theoretical analysis for structural perturbations (Theorem 5.4.2). To evaluate the effectiveness
of our approach in handling structural attacks, we conducted an empirical comparison to five baseline
defense methods that focus on structural perturbations, GNN-Jaccard (Wu et al., 2019b), RobustGCN (Zhu
et al., 2019), GNN-SVD (Entezari et al., 2020), GNNGuard (Zhang and Zitnik, 2020) and the Parseval
Regularization(ParsevalR) (Cisse et al., 2017). We considered three structural adversarial attacks with a
perturbation budget 0.1|E |: “Mettack” (with the ‘Meta-Self’ strategy) (Zügner and Günnemann, 2019), “Dice”
(Zügner and Günnemann, 2019) and “PGD” (Xu et al., 2019a). The average node classification accuracies
of the considered models under attack are presented in Table 5.2. Our GCORN shows remarkable defense
capabilities against these attacks, outperforming other methods in 8 of 12 experiments. This result highlights
the effectiveness of our proposed defense mechanism in enhancing the robustness of the underlying GCNs
against structural perturbations while providing theoretical guarantees which are absent for the considered
baselines.
5.7 C O N C L U S I O N 65

Table 5.2 – Attacked classification accuracy (± standard deviation) of the models on different benchmark node
classification datasets after the structural attacks.

Attack Dataset GCN GCN-Jaccard RGCN GNN-SVD GNN-Guard ParsevalR GCORN

Cora 73.0 ± 0.7 75.4 ± 1.8 69.2 ± 0.3 73.6 ± 0.9 74.4 ± 0.8 71.9 ± 0.7 77.3 ± 0.5
CiteSeer 63.2 ± 0.9 69.5 ± 1.9 68.9 ± 0.6 65.8 ± 0.6 68.8 ± 1.5 68.3 ± 0.8 73.7 ± 0.3
Mettack
PubMed 60.7 ± 0.7 62.9 ± 1.8 65.1 ± 0.4 82.1 ± 0.8 84.8 ± 0.3 69.5 ± 1.1 71.8 ± 0.4
CoraML 73.1 ± 0.6 75.4 ± 0.4 77.1 ± 1.1 71.3 ± 1.0 76.5 ± 0.7 76.9 ± 1.3 79.2 ± 0.6

Cora 76.7 ± 0.9 78.3 ± 1.1 72.0 ± 0.3 71.6 ± 0.4 75.0 ± 2.0 78.4 ± 1.2 79.9 ± 0.4
CiteSeer 67.8 ± 0.8 70.9 ± 1.0 62.2 ± 1.8 60.3 ± 2.4 68.9 ± 2.2 70.6 ± 1.0 73.1 ± 0.5
PGD
PubMed 75.3 ± 1.6 73.8 ± 1.3 78.6 ± 0.4 81.9 ± 0.4 84.3 ± 0.4 77.3 ± 0.7 77.4 ± 0.4
CoraML 76.9 ± 1.2 75.0 ± 2.4 77.5 ± 0.3 73.1 ± 0.5 75.5 ± 0.8 81.3 ± 0.4 84.1 ± 0.2

Cora 74.9 ± 0.8 76.9 ± 0.9 79.6 ± 0.3 72.2 ± 1.4 75.6 ± 1.1 79.7 ± 0.8 78.9 ± 0.4
CiteSeer 64.1 ± 0.5 66.0 ± 0.6 68.7 ± 0.5 62.6 ± 1.2 65.5 ± 1.1 68.9 ± 0.4 74.6 ± 0.4
DICE
PubMed 79.4 ± 0.4 78.3 ± 0.2 79.8 ± 0.4 76.6 ± 0.5 77.8 ± 0.7 79.2 ± 0.3 78.1 ± 0.6
CoraML 78.3 ± 0.6 77.5 ± 0.3 80.1 ± 0.4 58.7 ± 0.4 77.5 ± 0.2 80.5 ± 1.3 81.1 ± 0.8

5.7 CONCLUSION

We have proposed GCORN, an adaptation of the GCN, which is robust against feature-based adversarial
attacks. The approach utilizes the orthonormalization of weight matrices to control the robustness of GCNs
via an upper bound we derived. Additionally, we have proposed a probabilistic evaluation method for the
robustness of GNNs based on the introduced definition of Expected Adversarial Robustness. Experimental
results comparing our GCORN model to the standard GCN and existing defense methods show the superior
performance of GCORN on different real-world datasets.
R E T H I N K I N G R O B U S T N E S S I N G R A P H N E U R A L N E T W O R K S : A P O S T- H O C
APPROACH WITH CONDITIONAL RANDOM FIELDS
6
Neural Networks (GNNs), which are nowadays the benchmark approach in graph repre-

G
RAPH
sentation learning, have been shown to be vulnerable to adversarial attacks, raising concerns
about their real-world applicability. While existing defense techniques primarily concentrate on
the training phase of GNNs, involving adjustments to message passing architectures or pre-processing
methods, there is a noticeable gap in methods focusing on increasing robustness during inference. In
this context, this study introduces RobustCRF, a post-hoc approach aiming to enhance the robustness
of GNNs at the inference stage. Our proposed method, founded on statistical relational learning using a
Conditional Random Field, is model-agnostic and does not require prior knowledge about the underlying
model architecture. We validate the efficacy of this approach across various models, leveraging benchmark
node classification datasets.

6.1 INTRODUCTION

Deep Neural Networks (DNNs) have demonstrated exceptional performance across various domains,
including image recognition, language modeling, and speech recognition (Chen et al., 2019; Minaee et al.,
2024; Shim, Choi, and Sung, 2021). The growing interest in handling irregular and unstructured data,
particularly in domains like bioinformatics, has drawn significant attention to graph-based representations.
Graphs have emerged as the preferred format for representing such irregular data due to their ability to
capture interactions between elements, whether individuals in a social network or interactions between
atoms. In response to this need, Graph Neural Networks (GNNs) (Kipf and Welling, 2017a; Velickovic
et al., 2018; Xu et al., 2019c) have been proposed as an extension of DNNs tailored to graph-structured
data. GNNs excel in the generation of meaningful representations for individual nodes by leveraging both
a graph’s structural information and its associated features. This approach has demonstrated significant
success in addressing challenging applications in protein function prediction (Gilmer et al., 2017), materials
modeling (Duval et al., 2023b), time-varying data reconstruction, and recommendation systems (Wu et al.,
2019c).
Alongside their achievements, these deep learning-based approaches have exhibited vulnerability to
various data alterations, including noisy, incomplete, or out-of-distribution examples (Günnemann, 2022).
While such alterations may naturally occur in data collection, adversaries can deliberately craft and introduce
them, resulting in adversarial attacks. These attacks manifest as imperceptible modifications to the input
that can deceive and disrupt the classifier. In these attacks, the adversary’s objective is to introduce subtle
noise in the features or manipulate some edges in the graph structure to alter the initial prediction made on
the input graph. Depending on the attacker’s goals and level of knowledge, different attack settings can
be considered. For instance, poisoning attacks (Zügner and Günnemann, 2019) involve manipulating the
training data to introduce malicious data points, while evasion attacks (Dai et al., 2018) focus on attacking
the model during the inference phase without further model adaptation.
Given the susceptibility of GNNs to adversarial attacks, their practical utility is constrained. Therefore,
it becomes imperative to study and enhance their robustness. Various defense strategies have been
proposed to mitigate this vulnerability, including pre-processing the input graph (Entezari et al., 2020;
Wu et al., 2019b), edge pruning (Zhang and Zitnik, 2020), and adapting the message passing scheme

67
68 A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

(Abbahaddou et al., 2024a; Liu et al., 2021b; Zhu et al., 2019). Most of these defense methods operate
by modifying the underlying model architecture or the training procedure, i. e. focusing on the training
stage of GNNs and, therefore, limiting their applicability. Additionally, modifying the model architecture
during training poses the risk of degrading accuracy on clean, non-attacked graphs, and usually, these
modifications are not adapted for all possible architectures.
In light of these challenges, our work introduces a post-hoc defense mechanism, denoted as RobustCRF,
aimed at bolstering the robustness of GNNs during the inference phase using statistical relational learning.
RobustCRF is model-agnostic, requiring no prior knowledge of the underlying model, and is adaptable to
various architectural designs, providing flexibility and applicability across diverse domains. Central to our
approach is the assumption that neighboring points in the input manifold, accounting for graph isomorphism,
should yield similar predictions in the output manifold. Based on our robustness assumption, we employ a
Conditional Random Field (CRF) (Lafferty, McCallum, and Pereira, 2001) to adapt and edit the model’s
output to preserve the similarity relationship between the input and output manifold.
We start by introducing our CRF-based post-hoc robustness enhancement model. Recognizing the
potential complexity associated with this task and aiming to deliver a robustness technique that remains
computationally affordable, we study a sampling strategy for both the discrete space of the graph structure
and the continuous space of node features. We finally proceed with an empirical analysis to assess
the effectiveness of our proposed technique across various models, conducting also a comprehensive
examination of the parameters involved. In summary, our contributions can be outlined as follows. (1) Model-
agnostic robustness enhancement: We present RobustCRF, a post-hoc approach designed to enhance the
robustness of underlying GNNs, without any assumptions about the model’s architecture. (2) Theoretical
underpinnings and complexity reduction: We conduct a comprehensive theoretical analysis of our proposed
approach and enhance our general architecture through the incorporation of sampling techniques, thereby
mitigating the complexity associated with the underlying model. (3) Empirical Evaluation and Analysis:
We evaluate and compare our RobustCRF model to a wide array of baseline models on several node
classification benchmark datasets and observe convincing performance of our RobustCRF approach.

6.2 R E L AT E D W O R K

Attacking GNNs. A multitude of poisoning and evasion adversarial attacks targeting GNNs models
has surged lately (Günnemann, 2022; Zügner, Akbarnejad, and Günnemann, 2018). Gradient-based
techniques (Xu et al., 2019a), such as Proximal Gradient Descent (PGD), have been employed to tackle
the adversarial aim, framing it as an optimization task that seeks the closest adversarial example to
the input while satisfying the adversarial objective. Building upon this foundation, Mettack (Zügner and
Günnemann, 2019) extends the approach by expressing the problem as a bi-level optimization task and
harnessing meta-gradients for its solution. Taking a different perspective, Nettack (Zügner, Akbarnejad,
and Günnemann, 2018) introduces a targeted poisoning attack strategy, encompassing both structural
and node feature perturbations, employing a greedy optimization algorithm to minimize an attack loss with
respect to a surrogate model. Diverging from these classical search problems, Dai et al. (2018) approach
the adversarial search task using Reinforcement Learning techniques.
Defending GNNs. Different defense strategies have been proposed to counter the previously presented
attacks on GNNs. Some of these defense methods employ low-rank approximation of the adjacency
matrix to filter out noise (Alchihabi, En, and Guo, 2023; Entezari et al., 2020), while similar pre-processing
techniques are used to identify potential edge manipulations (Deng et al., 2022; Wu et al., 2019b).
Additionally, methods like edge pruning (Zhang and Zitnik, 2020) and transfer learning (Tang et al., 2020a,b;
Yang, Liu, and Mirzasoleiman, 2022) have been used to mitigate the impact of poisoning attacks. Notably,
most research efforts have predominantly focused on addressing structural perturbations, with fewer
6.2 R E L AT E D W O R K 69

strategies developed to counter attacks targeting node features. For instance, in Seddik et al. (2022),
the inclusion of node feature kernels in the message passing operator was proposed to increase GCN
robustness. RobustGCN (Zhu et al., 2019) uses Gaussian distributions as hidden node representations in
each convolutional layer, effectively mitigating the influence of both structural and feature-based adversarial
attacks. Finally, in GCORN (Abbahaddou et al., 2024a), an adaptation of the message passing scheme
has been proposed by using orthonormal weight matrices to counter the effect of node feature-based
adversarial attacks. Recent advances have further expanded the landscape of GNN defenses. For
instance, self-supervised learning has been explored as a robust defense mechanism against adversarial
perturbations (Yang et al., 2023). These approaches continue to push the boundaries of GNN security by
addressing both structural and feature-based vulnerabilities.
The majority of the previously discussed methods intervene during the model’s training phase, necessi-
tating modifications to the underlying architecture. However, this strategy exhibits certain limitations; it may
not be universally applicable across diverse architectural designs, and it can potentially result in a loss of
accuracy when applied to the clean graph. Furthermore, these methods may not be suitable for scenarios
where users prefer to employ pre-trained models, as they necessitate model retraining. Consequently,
the need for proposing and crafting post-hoc robustness enhancements becomes increasingly apparent.
Addressing this gap in the literature, our work aims to contribute to this essential area. One commonly
employed post-hoc approach is randomized smoothing (Bojchevski, Klicpera, and Günnemann, 2020;
Carmon et al., 2019), which involves injecting noise into the inputs at various stages and subsequently
utilizing majority voting to determine the final prediction. Note that randomized smoothing has actually
been initially borrowed from the optimization community (Duchi, Bartlett, and Wainwright, 2012). Despite
its popularity, randomized smoothing exhibits several limitations, including suffering from the “shrinkage
phenomenon” where decision regions shrink or drift as the variance of the smoothing distribution increases
(Mohapatra et al., 2021). Other works also identified that the smoothed classifier is more-constant than the
original model, i. e. it forces the classification to remain invariant over a large input space, resulting in a
drop in the accuracy (Anderson and Sojoudi, 2022; Krishnan et al., 2020; Wang et al., 2021a).
Conditional Random Fields. A probabilistic graphical model (PGM) is a graph framework that compactly
models the joint probability distributions P and dependence relations over a set of random variables
Ye = {Y em } represented in a graph. The two most common classes of PGMs are Bayesian Networks
e1 , . . . , Y
(BNs) and Markov Random Fields (MRFs) (Clifford, 1990; Heckerman, 2008). The core of the BN
representation is a directed acyclic graph (DAG). In the DAG representation of a BN, each node in the DAG
corresponds to a random variable and directed edges capture dependency relationships between these
random variables. The direction of the edges determines the influence of one random variable on another.
Similarly, MRFs are also used to describe dependencies between random variables in a graph. However,
MRFs use undirected instead of directed edges and permit cycles. An important assumption of MRFs is
the Markov property, i.e., for each pair of nodes ( a, b) that are not directly connected, i. e. eab ∈ / EMRF ,
node a is independent of node b conditioned on a’s neighbors,

∀ a ∈ V MRF , Y
ea ⊥ Y
eV \N (a) | N ( a).

An important special case of MRFs arises when they are applied to model a conditional probability
e | Y, V MRF , EMRF ), where Y = {Y1 , . . . , Ym } are additional observed node features. These
distribution P(Y
types of graphs are called Conditioned Random Fields (CRFs) (Sutton, McCallum, et al., 2012; Wallach
70 A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

Input Manifold Conditional Random Field (CRF)

Output Manifold
ons
predicti
Observables

CRF Dependencies

Figure 6.1 – Illustration of our RobustCRF approach. We use the input graph manifold to generate n the structure of
o
the CRF, i.e., V CRF , ECRF . We use the GNN’s predictions to generate the observables Ya : a ∈ V CRF ,
we then run the CRF inference to generate the new GNN’s predictions {Y ea : a ∈ V CRF }.

e and observation Y, with the


et al., 2004). Formally, a CRF is Markov network over random variables Y
conditional distribution defined as follows:

e | Y, V CRF , ECRF ) = 1
P (Y ×
Z (Y, V CRF , ECRF )
 
 
exp − ∑ ea ) −
ϕ a (Y ∑ ϕab (Y eb )
ea , Y ,
a∈V CRF ( a,b)∈ ECRF
 

where Z (Y, V CRF , ECRF ) is the partition function, and ϕa (Y


ea ), ϕab (Y eb ) are potential functions con-
ea , Y
tributed by each node a and edge ( a, b). They are usually defined as simple linear functions or learned by
simple regression on the features, e.g., with logistic regression.

6.3 PROPOSED METHOD: ROBUSTCRF

In this section, we first formally define adversarial attacks on models processing attributed graphs and
outline the common rationale followed in different robustness definitions. We then demonstrate how this
rationale can be represented via a CRF. We further illustrate how we can use the mean field approximation
and a sampling approach to fit a CRF to draw more robust inference from GNNs.

6.3.1 Adversarial Attacks

Let us consider our graph space (A, ∥·∥A ), feature space (X, ∥·∥X ), and the label space (Y, ∥·∥Y )
to be measurable spaces. Given a GNN f : (A, X) → Y, an input data point [A, X] ∈ A × X and
its corresponding prediction y ∈ Y, i.e., f (A, X) = y, the goal of an adversarial attack is to produce a
perturbed graph [A,
e Xe ] slightly different from the original graph with its predicted class being different from
the predicted class of [A, X]. This could be formulated as finding a [A, e ] with f (A,
e X e Xe ) ̸= y subject to
d([A, X], [A,
e X]) < ϵ, with d being some distance function between the original and perturbed graphs. This
e
could be a distance taking into account both the graph structure, in terms of the adjacency matrix, and the
corresponding node features, as defined in Equation 2.38.
6.3 P R O P O S E D M E T H O D : R O B U S T C R F 71

6.3.2 Motivation

There are different theoretical definitions of robustness (Cheng et al., 2021; Weng et al., 2018b) and
the great majority, if not all, rely on one assumption: If two inputs are adjacent in the input space, their
predictions should be adjacent in the output space. Many adversarial attack methods use this principle by
attempting to find small perturbations that cause significant changes in the model’s output, highlighting its
vulnerability (Goodfellow, Shlens, and Szegedy, 2015b; Wu et al., 2021). Additionally, there are also works
that use the neural networks distortion as a robustness metric (Carlini and Wagner, 2016; Cheng et al.,
2021; Weng et al., 2018b). Intuitively, a large distortion implies potentially poor adversarial robustness since
a small perturbation applied to these inputs will lead to significant changes in the output. Consequently,
most robustness metrics measure the extent to which the network’s output is changed when perturbing the
input data, indicating the network’s vulnerability to adversarial attacks. To respect this common assumption,
we construct a CRF, where its node set V CRF represents the set of possible GNN inputs A × X, while its
edge set ECRF represents the set of pair inputs ([A, X], [A, e ]) such that [A,
e X e Xe ] belongs to the ball B of
radius r > 0 surrounding [A, X], as follows,
nh i o
B ([A, X] , r ) = A,e Xe : dα,β ([A, X], [A,
e Xe ]) ≤ r .

6.3.3 Modeling the Robustness Constraint with a CRF

Let Ya be the output prediction of the trained GNN f on the input a = [A, X] ∈ V CRF . The main goal of
using a CRF is to update the predictions Y = {Ya : a ∈ V CRF } into new predictions Ye = {Yea : a ∈ V CRF }
that respect the robustness assumption. To do so, we model the relation between the two predictions Y
and Ye using a CRF, maximizing the following conditional probability

e | Y, V CRF , ECRF ) = 1 exp
P (Y − ∑ ϕ a (Y
ea , Ya )
Z a∈V CRF

− ∑ ϕab (Y eb ) , (6.1)
ea , Y
( a,b)∈ ECRF

where Z is the partition function, ϕa (Y


ea , Ya ) and ϕab (Y eb ) are potential functions contributed by each CRF
ea , Y
node a and each CRF edge ( a, b) and usually defined as simple linear functions or learned by simple
regression on the features (e.g., logistic regression). The potential functions ϕa and ϕab can be either fixed
or trainable to optimize a specific objective. In this paper, we define the potential functions as follows:
ea − Ya ∥22 ,
ea , Ya ) = σ ∥Y
ϕ a (Y
ϕab (Y eb ) = (1 − σ ) gab ∥Y
ea , Y eb ∥22 ,
ea − Y

where gab denotes a chosen similarity term of two inputs a = [A, X] and b = [A, e Xe ], e.g., the cosine
similarity between their feature matrices. The parameter σ is used to adjust the importance of the two
potential functions. Representing GNN inputs in the CRF can be seen as a constrained problem to enhance
the robustness: Finding a new prediction Y ea that satisfies our robustness assumption and, at the same
time, stays as close as possible to the original prediction of the GNN, Ya = f ( a) where a = [A, X] is a GNN
input. In Figure 6.1, we illustrate the main idea of the proposed RobustCRF model.
Now that we have defined the CRF and its potential functions, generating the new prediction {Y ep : p ∈
V CRF } is intractable for two reasons. First, the partition function Z is usually intractable. Second, the CRF
distribution P represents a potentially infinite collection of CRF edges ECRF . We need, therefore, to derive
72 A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

a tractable algorithm to generate the smoothed prediction Y.


e In what follows, we show how to overcome
these two challenges.

6.3.4 Mean Field Approximation

We aim to derive the most likely Y


e from the initial distribution P defined in (6.1), as,

e∗ = arg max P(Y


Y e | Y, V CRF , ECRF ). (6.2)
e ∈Y
Y

Since the inference is intractable, we use a variational inference method where we propose a family of
densities D and find a member Q∗ ∈ D which is close to the posterior P(Y e | Y, V CRF , ECRF ). Thus, we
approximate the initial task in (6.2) with a new objective,

e∗ ∗ e
 Y = arg max Q (Y ),


e ∈Y
Y (6.3)

 Q ∗ = arg min KL ( Q | P ) ,

Q∈D

where KL (· | ·) denotes the Kullback-Leibler divergence. The goal is to find the distribution Q∗ within the
family D , which is the closest to the initial distribution P. In this work, we used the mean-field approximation
that enforces full independence among all latent variables (Blei, Kucukelbir, and McAuliffe, 2017). Mean
field approximation is a powerful technique for simplifying complex probabilistic models, widely used in
various fields of machine learning and statistics (Andrews and Baguley, 2017). Thus, the variational
distribution over the latent variables factorizes as,

∀ Q ∈ D , Q (Y
e) = ∏ Q a (Y
ea ). (6.4)
a∈V CRF

The exact formula of the optimal surrogate distribution Q can be obtained using Lemma 6.3.1.

Lemma 6.3.1. By solving the system of (6.3), we can get the optimal distribution Q∗ for each a ∈ V CRF as
follows, n h  io
ea ) ∝ exp E−a log P Y
Q a (Y e | Y, V CRF , ECRF , (6.5)

where E−a denotes the expectation taken over the random variables Y
e−a , corresponding to all nodes except
a,

The proof of Lemma 6.3.1 is provided in Appendix d.2. We will use Coordinate Ascent Inference (CAI),
iteratively optimizing each variational distribution and holding the others fixed. The CAI iteratively updates
each Q a (Yea ). The evidence lower bound (ELBO) converges to a local minimum. Using (6.1) and (6.5), we
get the optimal surrogate distribution as follows,
( )
ea ) ∝ exp σ ∥Y
Q a (Y ea − Ya ∥22 + (1 − σ ) ∑ gab ∥Y eb ∥22 .
ea − Y
b s.t. ( a,b)∈ ECRF

Thus, for each GNN input a ∈ V CRF , Q a (Y


ea ) is a Gaussian distribution that reaches the highest probability
at its expectation,
σYa + (1 − σ ) ∑ gab Y
eb
b s.t. ( a,b)∈ ECRF
arg max Q a (Y
ea ) = .
ea ∈Y
Y
σ + (1 − σ ) ∑ gab
b s.t. ( a,b)∈ ECRF
6.3 P R O P O S E D M E T H O D : R O B U S T C R F 73

Using the CAI algorithm, we can thus utilize the following update rule:

σYa + (1 − σ ) ∑ ek
gab Yb
b s.t. ( a,b)∈ ECRF
eak+1 =
Y . (6.6)
σ + (1 − σ ) ∑ gab
b s.t. ( a,b)∈ ECRF

6.3.5 Reducing the Size of the CRF

The number of possible inputs is usually very large, or infinite, for discrete data and infinite for continuous
data. This could make the inference intractable due to the potential large size of ECRF . Therefore, instead
of considering all the possible CRF neighbors of an input a, i.e., all inputs b ∈ V CRF such that dα,β ( a, b) ≤ r,
we can consider only a subset of L neighbors by randomly sampling from the CRF neighbors of a. The
update rule in (6.6) can then be expressed as,

σYa + (1 − σ ) ∑ ek
gab Yb
b∈U L ( a)
eak+1 =
Y , (6.7)
σ + (1 − σ ) ∑ gab
b∈U L ( a)

where U L ( a) denotes a set of L randomly sampled CRF neighbors of a graph [A, X].

Remark 6.3.2. If in (6.7) we set the value of σ to 0, the number of iterations to 1, and all the similarity coef-
ficients gab to 1, this scheme corresponds to the standard randomized smoothing with uniform distribution.
In this setting, we do not take into account the original classification task, which causes a huge drop in the
clean accuracy. Therefore, RobustCRF is a generalization of the randomized smoothing that gives a better
trade-off between accuracy and robustness.

Below, we elaborate on how to uniformly sample a neighbor graph b surrounding a in the structural
distances, i.e., (α, β) = (1, 0).
2
Let A = {0, 1}n be the adjacency matrix space, where n is the number of nodes. A is a finite-
dimensional compact normed vector space, so all the norms are equivalent. Thus, all the induced L p
distances are equivalent. Without loss of generality, we can consider the L1 loss, which exactly corresponds
to the Hamming distance,
d1 ([A, X], [A,
e Xe ]) = ∑ |Ai,j − A
e i,j |,
i≤ j

where A, Ae correspond to the adjacency matrices of graphs G, G,


e respectively. Since we consider undi-
n ( n +1)
rected graphs, the Hamming distance takes only discrete values in {0, . . . , 2 }. In Lemma 6.3.3, we
provide a lower bound for the size of CRF neighbors N CRF ( a) = {b ∈ V CRF : ( a, b) ∈ ECRF } for any
GNN input a ∈ V CRF .
n ( n +1)
Lemma 6.3.3. For any integer r in {0, . . . , 2 }, the number of CRF neighboors N CRF ( a) for any
a ∈ V CRF , i.e., the set of graphs [A,
e Xe ] with a Hamming distance smaller or equal than r, for each a ∈ V CRF ,
we have the following lower bound:

2 H (ϵ)n(n+1)/2
p ≤ N CRF ( a) , (6.8)
4n(n + 1)ϵ(1 − ϵ)
2r
where 0 ≤ ϵ = n ( n +1)
≤ 1 and H (·) is the binary entropy function, i.e., H (ϵ) = −ϵ log2 (ϵ) − (1 −
ϵ) log2 (1 − ϵ).
74 A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

Table 6.1 – Attacked classification accuracy (± standard deviation) of the GCN, the baselines and the proposed
RobustCRF on different benchmark node classification datasets after the features-based attack. ⃝
1 Test
accuracy on the original features; ⃝
2 Test accuracy on the perturbed features.

Attack Model Cora CiteSeer PubMed CS Texas

GCN 80.66±0.41 70.37±0.53 78.16±0.67 89.15±2.06 51.35±18.53


RGCN 77.64±0.52 69.88±0.47 75.58±0.65 92.05±0.72 51.62±13.91
Clean

1 GCORN 77.83±2.33 71.68±1.54 76.03±1.29 88.94±1.86 59.73±4.90
NoisyGCN 81.04±0.74 70.36±0.79 78.13±0.53 91.47±0.92 48.91±20.02
RobustCRF 80.63±0.38 70.30±0.43 78.20±0.24 88.16±3.41 61.08±5.01

GCN 77.88±0.90 66.65±1.00 73.60±0.75 88.92±2.04 46.49±15.75


RGCN 67.61±0.80 59.76±1.01 61.93±1.18 90.74±1.08 43.51±10.22
Random
GCORN 76.28±1.96 67.82±2.18 72.35±1.43 88.31±2.10 60.00±4.95
(ψ = 0.5)
NoisyGCN 78.59±1.09 66.83±0.98 73.60±0.58 91.04±0.85 48.91±19.29
RobustCRF 78.28±0.68 68.23±0.57 74.37±0.47 88.30±3.25 57.30±4.32

2
GCN 76.38±0.72 67.57±0.77 74.86±0.65 86.90±1.91 52.97±19.19
RGCN 68.45±0.97 64.63±0.82 73.35±0.81 90.76±0.68 60.81±10.27
PGD GCORN 73.32±2.19 69.05±2.50 74.49±1.13 87.07±2.96 59.73±4.90
NoisyGCN 76.29±1.69 67.09±1.50 75.04±0.53 88.79±0.85 52.97±19.11
RobustCRF 76.41±0.70 67.90±0.63 75.17±0.91 85.60±2.75 62.16±4.83

In Appendix d.1, we present the proof of Lemma 6.3.3 and we empirically analyze the asymptotic
behavior of this lower bound, motivating, thus, the need for sampling strategies to reduce the size of the
CRF.
To sample from B ([A, X] , r ), we use a stratified sampling strategy. First, we sample a distance value
d from {0, 1, . . . , n2 }. To do so, we partition B ([A, X ] , r ) with respect to their distance to the original
adjacency matrix A, 
 Sd = { A

 e ∈ A : d1 (A, A e ) = d},
B([A, X], r ) = d≤r Sd ,
S

 ∀d ̸= d′ , S ∩ S ′ = ∅.

d d

To each distance value d, we assign the portion of graphs covered by Sd (A) in A as


 
r 1
∀d ∈ {0, . . . , r }, p(d) = ,
d R
where R = ∑rd=0 (dr ) = 2r . Second, we uniformly select d positions in the adjacency matrix to be modified
and change the element of the d positions by changing the value Ai,j to 1 − Ai,j . To uniformly sample
neighbor graphs when dealing with feature-based distances for graphs, we used the sampling strategy of
Abbahaddou et al. (2024a).
RobustCRF is an attack-independent and model-agnostic robustness approach based on uniform
sampling without requiring any training. Therefore, it can be used to enhance GNNs’ robustness against
6.3 P R O P O S E D M E T H O D : R O B U S T C R F 75

Table 6.2 – Attacked classification accuracy (± standard deviation) of the baselines when combined with the pro-
posed RobustCRF on different benchmark node classification datasets after the features-based attack
application.

Attack Model Cora CiteSeer PubMed Texas

RGCN 77.64±0.52 69.88±0.47 75.58±0.65 51.62±13.91


RGCN w/ RobustCRF 77.70±0.46 69.84±0.39 75.50±0.60 52.16±14.20
GCORN 77.83±2.33 71.68±1.54 76.03±1.29 59.73±4.90
Clean
GCORN w/ RobustCRF 78.50±1.17 71.72±1.46 76.13±1.08 59.77±4.68
NoisyGCN 81.04±0.74 70.36±0.79 78.13±0.53 48.91±20.02
NoisyGCN w/ RobustCRF 81.07±0.70 70.20±0.85 78.90±0.46 49.18±19.59

RGCN 67.61±0.80 59.76±1.01 61.93±1.18 43.51±10.22


RGCN w/ RobustCRF 69.08±0.73 60.04±1.01 63.05±0.88 45.05±9.99
Random GCORN 76.28±1.96 67.82±2.18 72.35±1.43 60.00±4.95
(ψ = 0.5) GCORN w/ RobustCRF 77.21±1.17 68.94±3.00 72.29±1.29 59.18±3.71
NoisyGCN 78.59±1.09 66.83±0.98 73.60±0.58 48.91±19.29
NoisyGCN w/ RobustCRF 81.07±0.99 67.04±1.38 74.18±0.98 45.94±14.54

RGCN 68.45±0.97 64.63±0.82 73.35±0.81 60.81±10.27


RGCN w/ RobustCRF 68.46±0.93 64.59±0.84 73.47±0.72 58.10±11.09
GCORN 73.32±2.19 69.05±2.50 74.49±1.13 59.73±4.90
PGD
GCORN w/ RobustCRF 73.65±1.55 69.09±2.57 74.59±0.90 60.27±4.68
NoisyGCN 76.29±1.69 67.09±1.50 75.04±0.53 52.97±19.11
NoisyGCN w/ RobustCRF 76.48±1.65 67.21±1.35 75.39±0.45 53.51±19.07

unknown attack distributions. In Section 6.4, we will experimentally validate this theoretical insight for the
node classification task, demonstrating that RobustCRF achieves a good trade-off between robustness
and clean accuracy, i. e. the model’s initial performance on a clean un-attacked dataset. Moreover, the
approach of our work and these baselines are fundamentally different, since our RobustCRF is post-hoc,
we can also use the baselines in combination with our proposed RobustCRF approach to make even more
robust predictions. We report the results of this experiment in Table 6.2.
An advantage of RobustCRF is that it allows flexibility in defining the CRF structure based on different
criteria. For instance, we can construct the CRF by considering the worst-case scenario, where the
CRF neighbors correspond to the most adversarial perturbations within a given threat model. However,
instead of evaluating the model’s behavior only under specific adversarial attacks, our method examines its
performance more generally within a defined neighborhood. This perspective leans towards a concept of
average robustness, expanding on the conventional worst-case based adversarial robustness that is often
emphasized in adversarial studies (Abbahaddou et al., 2024a). A similar average robustness concept was
studied and showed to be appropriate for computer vision (Rice et al., 2021b).
76 A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

Table 6.3 – Attacked classification accuracy (± standard deviation) of the GCN, the baselines and the proposed
RobustCRF on different benchmark node classification datasets after the structural attack application.

1 Test accuracy on the original structure, ⃝
2 Test accuracy on the perturbed structure.

Attack Model Cora CoraML CiteSeer PolBlogs

GCN 83.42±1.00 85.60±0.40 70.66±1.18 95.16±0.64


RGCN 83.46±0.53 85.61±0.61 72.18±0.97 95.32±0.76
GCNGuard 83.72±0.67 85.54±0.42 73.18±2.36 95.07±0.51
Clean

1 GCNSVD 77.96±0.61 81.29±0.51 68.16±1.15 93.80±0.73
GCNJaccard 82.20±0.67 84.85±0.39 73.57±1.21 51.81±1.49
GOOD-AT 83.43±0.11 84.87±0.15 72.80±0.45 94.85±0.52
RobustCRF 83.52±0.04 85.69±0.09 72.16±0.30 95.40±0.24

GCN 81.87±0.73 83.34±0.60 71.76±1.06 87.14±0.86


RGCN 81.27±0.71 83.89±0.51 69.45±0.92 87.35±0.76
GCNGuard 81.63±0.74 83.72±0.49 71.87±1.19 86.98±1.26
Dice

2 GCNSVD 75.62±0.61 79.13±1.01 66.10±1.29 88.50±0.97
GCNJaccard 80.67±0.66 82.88±0.58 72.32±1.10 51.81±1.49
GOOD-AT 82.21±0.56 84.16±0.36 71.43±0.28 90.93±0.38
RobustCRF 82.44±0.41 84.71±0.32 71.48±0.15 90.46±0.25

6.4 E X P E R I M E N TA L R E S U LT S

In this section, we shift from theoretical exploration to practical validation by evaluating the effectiveness
of RobustCRF on real-world benchmark datasets. Our primary experimental objective is to assess how
well the proposed method enhances the robustness of a trained GNN in a node classification task. Details
about the datasets used in our experiments and implementation details can be found in Appendix d.4.
Attacks. We evaluate RobustCRF via the Attack Success Rate (ASR), the percentage of attack attempts
that produce successful adversarial examples. For the feature-based attacks, we consider two main types:
(1) we first consider a random attack which consists of injecting Gaussian noise N (0, ψI) to the features
with scaling parameter ψ = 0.5; (2) we have additionally used the white-box Proximal Gradient Descent
(Xu et al., 2019a), which is a gradient-based approach to the adversarial optimization task for which we set
the perturbation rate to 15%. For the structural perturbations, we evaluated RobustCRF using the “Dice”
adversarial attack in a black-box setting, where we consider a surrogate model (Zügner and Günnemann,
2019). For this setting, we used an attack budget of 10% (the ratio of perturbed edges).
Baseline Models. When dealing with feature-based attacks, we compare RobustCRF with the vanilla
GCN (Kipf and Welling, 2017a), the feature-based defense method RobustGCN (RGCN) (Zhu et al., 2019),
NoisyGNN (Ennadir et al., 2024), and GCORN (Abbahaddou et al., 2024a). For the structural attacks,
we again included the RGCN and other baselines such as GNN-Jaccard (Wu et al., 2019b), GNN-SVD
(Entezari et al., 2020), GNNGuard (Zhang and Zitnik, 2020), and GOOD-AT (Li et al., 2024a). For all the
models, we used the same number of layers T = 2, with a hidden dimension of 16. The models were
trained using the cross-entropy loss function with the Adam optimizer (Kingma and Ba, 2015a), the number
of epochs Nepochs = 300, and learning rate 0.01 were kept similar for the different approaches across all
experiments. To reduce the impact of random initialization, we repeated each experiment 10 times and
6.5 C O N C L U S I O N 77

used the train/validation/test splits provided with the datasets when evaluating against the feature-based
attacks, c.f. Table 6.1. When evaluating against the structural attacks, c.f. Table 6.3, we used the split
strategy of (Zügner, Akbarnejad, and Günnemann, 2018), i.e., we select the largest connected components
of the graph and use 10%/10%/80% nodes for training/validation/test.
Worst-Case Adversarial Evaluation. We now analyze the results of RobustCRF for the node classi-
fication task. Additionally, we have compared our approach to the baseline methods on OGBN-Arxiv, a
large dataset, as detailed in Appendix d.6. We report all the results for the feature and structure-based
adversarial attacks, respectively, in Tables 6.1, 6.3, and d.5. The results demonstrate that the performance
of the GCN model is significantly impacted when subject to adversarial attacks of varying strategies. In
contrast, the proposed RobustCRF approach shows robust performance against these attacks compared
to other baseline models. Furthermore, in contrast to some other benchmarks, RobustCRF offers a
good balance between robustness and clean accuracy. Specifically, RobustCRF effectively enhances the
robustness against adversarial attacks while maintaining high accuracy on non-attacked, clean datasets.
This latter point makes RobustCRF particularly advantageous, as it enhances the model’s defenses without
compromising its performance on downstream tasks.
Enhancing Baseline Defenses with RobustCRF. One key advantage of RobustCRF is its ability to
integrate with other defense methods, enhancing their overall robustness. In Table 6.2, we report the
performance of the baselines when combined with RobustCRF. As noticed, for most of the cases, we
further enhance the robustness of the baselines when using RobustCRF. This demonstrates its adaptability
to different defense strategies, particularly those operating at the architectural level.
Time and Complexity. We analyze the computational complexity of RobustCRF, focusing on its inference
phase, which depends primarily on the number of iterations K and the number of sampled neighbors
L. Since RobustCRF applies iterative updates following a coordinate ascent inference (CAI) scheme, its
complexity is influenced by the size of the CRF and the number of updates required for convergence.
In Appendix d.5, we report the average time needed to compute the RobustCRF inference. The results
validate the intuitive fact that the inference time grows by increasing K and L. We recall that we need to
use the model LK times in the CRF inference. While increasing K and L can improve robustness, it also
increases computational cost. In practical settings, we observe that convergence is often reached within a
small number of iterations, making RobustCRF computationally feasible even for large graphs.

6.5 CONCLUSION

This work addresses the problem of adversarial defense at the inference stage. We propose a model-
agnostic, post-hoc approach using Conditional Random Fields (CRFs) to enhance the adversarial ro-
bustness of pre-trained models. Our method, RobustCRF, operates without requiring knowledge of the
underlying model and necessitates no post-training or architectural modifications. Extensive experiments
on multiple datasets demonstrate RobustCRF’s effectiveness in improving the robustness of Graph Neural
Networks (GNNs) against both structural and node-feature-based adversarial attacks, while maintaining a
balance between attacked and clean accuracy, typically preserving their performance on clean, un-attacked
datasets, which makes RobustCRF the best trade-off between robustness and clean accuracy.
Part IV

G E N E R A L I Z AT I O N O F G R A P H N E U R A L N E T W O R K S
I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N
7
Neural Networks (GNNs) have shown great promise in tasks like node and graph classification,

G
RAPH
but they often struggle to generalize, particularly to unseen or out-of-distribution (OOD) data.
These challenges are exacerbated when training data is limited in size or diversity. To address
these issues, we introduce a theoretical framework using Rademacher complexity to compute a regret
bound on the generalization error and then characterize the effect of data augmentation. This framework
informs the design of GRATIN, an efficient graph data augmentation algorithm leveraging the capability of
Gaussian Mixture Models (GMMs) to approximate any distribution. Our approach not only outperforms
existing augmentation techniques in terms of generalization but also offers improved time complexity,
making it highly suitable for real-world applications.

7.1 INTRODUCTION

Graphs are a fundamental and ubiquitous structure for modeling complex relationships and interactions.
In biology, graphs are employed to represent complex networks of protein interactions and in drug discovery
by modeling molecular relationships (Gaudelet et al., 2021; Jagtap et al., 2022). Similarly, social networks
capture relationships and community interactions (Malliaros and Vazirgiannis, 2013b; Newman, Watts,
and Strogatz, 2002; Zeng et al., 2022). To address the unique challenges posed by graph-structured
data, GNNs have been developed as a specialized class of neural networks designed to operate directly
on graphs. Unlike traditional neural networks that are optimized for grid-like data, such as images or
sequences, GNNs are engineered to process and learn from the relational information embedded in graph
structures. GNNs have demonstrated state-of-the-art performance across a range of graph representation
learning tasks such as node and graph classification, proving their effectiveness in various real-world
applications (Castro-Correa et al., 2024b; Chi et al., 2022; Corso et al., 2022; Duval et al., 2023b; Vignac
et al., 2023).
Despite their impressive capabilities, GNNs face significant challenges related to generalization, particu-
larly when handling unseen or out-of-distribution (OOD) data (Guo et al., 2024; Li et al., 2022). OOD graphs
are those that differ significantly from the training data in terms of graph structure, node features, or edge
types, making it difficult for GNNs to adapt and perform well on such data. This challenge is also faced
when GNNs are trained on small datasets, where the limited data diversity hampers the model’s ability
to generalize effectively. To address these challenges, the community has explored various strategies to
improve the robustness and generalization ability of GNNs (Abbahaddou et al., 2024a; Yang et al., 2022).
Generalization bounds for GNNs have been derived using various theoretical tools, such as the Vapnik-
Chervonenkis (VC) dimension (Garg, Jegelka, and Jaakkola, 2020; Pfaff et al., 2020) and Rademacher
complexity (Esser, Chennuru Vankadara, and Ghoshdastidar, 2021; Yin, Kannan, and Bartlett, 2019).
Furthermore, Liao, Urtasun, and Zemel (2020) were among the first to establish generalization bounds
for GNNs using the PAC-Bayesian approach. Neural Tangent Kernels have also been employed to study
the generalization properties of infinitely wide GNNs trained via gradient descent (Du et al., 2019; Huang
et al., 2024; Jacot, Gabriel, and Hongler, 2018). While most existing research has focused on the node
classification task, enhancing generalization in graph classification presents unique challenges. Techniques
to improve generalization in graph classification can be broadly categorized into architectural and dataset-
based strategies (Buffelli, Liò, and Vandin, 2022; Tang and Liu, 2023). On the dataset side, techniques

81
82 I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

like adversarial training and data augmentation play a significant role. More broadly, data augmentation
methods create synthetic or modified graph instances to enrich the training set, reducing overfitting and
enhancing the model’s adaptability to diverse graph structures. Data augmentation has shown its benefits
across different types of data structures such as images (Krizhevsky, Sutskever, and Hinton, 2012) and
time series (Aboussalah et al., 2023). For graph data structures, generating augmented versions of the
original graphs, such as by adding or removing nodes and edges or perturbing node features (Rong et al.,
2019; You et al., 2020), allows for the creation of a more varied training set. Inspired by the success of the
Mixup technique in computer vision (Dabouei et al., 2021; Hong, Choi, and Kim, 2021; Rebuffi et al., 2021),
additional methods such as G -Mixup and GeoMix have been developed to adapt Mixup for graph data (Han
et al., 2022; Ling et al., 2023). These techniques combine different graphs to create new, synthetic training
examples, further enriching the dataset and enhancing the GNN’s ability to generalize to new unseen graph
structures.
In this work, we introduce GRATIN, a graph augmentation technique based on Gaussian Mixture Models
(GMMs), which operates at the level of the final hidden representations. Specifically, guided by our
theoretical results, we apply the Expectation-Maximization (EM) algorithm to train a GMM on the graph
representations. We then use this GMM to generate new augmented graph representations through
sampling, enhancing the diversity of the training data.
The contributions of our work are as follows:
— Theoretical framework for generalization in GNNs. We introduce a theoretical framework that
allows us to rigorously analyze how graph data augmentation impacts the generalization GNNs. This
framework offers new insights into the underlying mechanisms that drive performance improvements
through augmentation.
— Efficient graph data augmentation via GMMs. We propose GRATIN, a fast and efficient graph data
augmentation technique leveraging GMMs. This approach enhances the diversity of training data
while maintaining computational simplicity, making it scalable for large graph datasets.
— Comprehensive theoretical analysis using influence functions. We perform an in-depth the-
oretical analysis of our augmentation strategy through the lens of influence functions, providing a
principled understanding of the approach’s impact on generalization performance.
— Empirical Validation. Through experiments on real-world datasets we confirm GRATIN to be a fast,
high-performing graph augmentation scheme in practice.

7.2 B A C K G R O U N D A N D R E L AT E D W O R K

Graph data augmentation has become essential to enhance the performance and robustness of GNNs.
Classical graph augmentation techniques focus on structural modifications to generate augmented graphs.
Key methods here include DropEdge, DropNode, and Subgraph sampling (Rong et al., 2019; You et al.,
2020). For instance, DropEdge randomly removes a subset of edges from the graph during training,
improving the model’s robustness to missing or noisy connections. Similarly, DropNode removes certain
nodes as well as their connections, assuming that the missing part of nodes will not affect the semantic
meaning, i.e., the structural and relational information of the original graph. Subgraph sampling, on the
other hand, samples a subgraph from the original graph using random walks to use as a training graph.
Beyond classical methods, recent advancements have explored more sophisticated augmentation
techniques, focusing on manipulating graph embeddings and leveraging the geometric properties of graphs.
Following the effectiveness of the Mixup technique in computer vision (Dabouei et al., 2021; Hong, Choi,
and Kim, 2021; Rebuffi et al., 2021), several works describe variations of the Mixup for graphs. For example,
the Manifold-Mixup model conducts a Mixup operation for graph classification in the embedding space.
7.3 G R AT I N : G A U S S I A N M I X T U R E M O D E L F O R G R A P H D ATA A U G M E N TAT I O N 83

This technique interpolates between graph-level embeddings after the READOUT function, blending different
graphs in the embedding space (Wang et al., 2021b). Similarly, G -Mixup (Han et al., 2022) uses graphons
to model the topological structures of each graph class and then interpolates the graphons of different
classes, subsequently generating synthetic graphs by sampling from mixed graphons across different
classes. It is important to note that G -Mixup operates under a significant assumption: graphs belonging
to the same class can be produced by a single graphon. Other advanced techniques include S-Mixup
method, which interpolates graphs by first determining node-level correspondences between a pair of
graphs (Ling et al., 2023), and FGW-Mixup, which adopts the Fused Gromov-Wasserstein barycenter to
compute mixup graphs but suffers from heavy computation time (Ma et al., 2024). Finally, GeoMix (Zeng
et al., 2024) leverages Gromov-Wasserstein geodesics to interpolate graphs more efficiently. By leveraging
these structural augmentation techniques, GNNs can better generalize to unseen graph structures.

7.3 G R AT I N : G A U S S I A N M I X T U R E M O D E L F O R G R A P H D ATA A U G M E N TAT I O N

In this section, we introduce the mathematical framework for graph data augmentation and its connection
to the generalization of GNNs. Then, we present our proposed model GRATIN, which is based on GMMs
for graph augmentation.

7.3.1 Formalism of Graph Data Augmentation

We focus on the task of graph classification, where the objective is to classify graphs into predefined
categories. Let D denote the distribution of graphs. Given a training set of graphs Dtrain = {(Gn , yn ) | n =
1, . . . , N }, Gn is the n-th graph and yn is its corresponding label belonging to a set {0, . . . , C }. Each graph
Gn is represented as a tuple (Vn , En , Xn ), where Vn denotes the set of nodes with cardinality pn = |Vn |,
En ⊆ Vn × Vn is the set of edges, and Xn ∈ R pn ×d is the node feature matrix of dimension d. The objective
is to train a GNN f (·, θ ) that can accurately predict the class labels for unseen graphs in the test set
Dtest = {Gntest | n = 1, . . . , Ntest }. The classical training approach involves minimizing the following loss
function,
L = ℓ( f (Gn , θ ), yn ), (7.1)
where ℓ denotes the cross-entropy loss function.
To improve the generalization performance of GNNs, we introduce a graph data augmentation strategy.
For each training graph Gn in the dataset, we generate M augmented graphs, denoted as {Gen,m , yen,m |
m = 1, . . . , M }, where M is the number of augmented graphs generated per training graph. These
augmented graphs are obtained using a graph augmentation generator Aλ , parameterized by λ as a
mapping Aλ : Gn , yn → Aλ (Gn , yn ) ∈ G × Y, where G denotes the space of all possible graphs, and Y is
the label space. The generator Aλ may be either deterministic or stochastic with a dependence on a prior
distribution P (λ). Examples of such augmentation strategies can be found in Appendix e.8.
We use the notation Genm ∼ Aλ to represent an augmented graph sampled from the augmentation strategy
Aλ , i.e., (Genm , yem
n ) ∼ Aλ (Gn , yn ). With the augmented data, the loss function is modified to account for
multiple augmented versions of each graph,
N
1 h i
Laug =
N ∑ EGe m
n ∼ Aλ
ℓ( f (Genm , θ ), yem
n) . (7.2)
n =1
84 I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

For simplicity, we denote the loss


h for the original i graph by ℓ( f (Gn , θ ), yn ) = ℓ(Gn , θ ) and the loss for
an augmented graph as EGem ∼ Aλ ℓ( f (Genm , θ ), yem
n) = ℓ
aug ( G
en , θ ). Via the law of large numbers, Laug is
n
empirically estimated,
N
1
Laug =
N ∑ ℓaug (Gen , θ )
n =1
N M
1

NM ∑ ∑ ℓ( f (Genm , θ ), yem
n ).
n =1 m =1

To understand the impact of data augmentation on the graph classification performance, we analyze the
effect of the augmentation strategy on the generalization risk EG∼D [ℓ(G , θ )]. More specifically, we want to
study the generalization error,

η = EG∼D [ℓ(G , θaug )] − EG∼D [ℓ(G , θ⋆ )] ,

where θaug and θ⋆ are the optimal GNN parameters for the augmented and non-augmented settings,
respectively,
aug
θ⋆ = arg min Lθ , θaug = arg min Lθ , (7.3)
θ θ
which can be estimated empirically as follows,
N
1
θ̂ = arg min
N ∑ ℓ(Gn , θ ),
θ n =1
N M
1
θ̂aug = arg min
NM ∑ ∑ ℓ(Genm , θ ).
θ n =1 m =1

By theoretically studying the generalization error η, we aim to quantify the effect of each augmentation
strategy on the overall classification performance, providing insights into the benefits and potential trade-offs
of data augmentation in graph-based learning tasks. In Theorem 7.3.1, we present a regret bound of the
generalization error using Rademacher complexity defined as follows Yin, Kannan, and Bartlett, 2019,
" #
1 N
R(ℓ) = Eϵn ∼ Pϵ sup ∑ ϵn ℓ(Gn , θ ) ,
θ ∈ Θ N n =1

where ϵn are independent Rademacher variables, taking values +1 or −1 with equal probability, Pϵ is
the Rademacher distribution, and Θ is the hypothesis class. Rademacher complexity is a fundamental
concept in statistical learning, which indicates how well a learned function will perform on unseen data
(Shalev-Shwartz and Ben-David, 2014). Intuitively, the Rademacher complexity measures the capacity of a
GNN to fit random noise Zhu, Gibson, and Rogers, 2009, where a lower Rademacher complexity indicates
better generalization.
Theorem 7.3.1. Let ℓ be a classification loss function with LLip as a Lipschitz constant and ℓ(·, ·) ∈ [0, 1].
Then, with a probability at least 1 − δ over the samples Dtrain , we have,

EG∼D ℓ(G , θ̂aug ) − EG∼D [ℓ(G , θ⋆ )] ≤ 2R(ℓaug )+


 

(7.4)
r
2 log(4/δ) h i
5 + 2LLip EG∼D ,G∼
e Aλ Ge − G .
N
Moreover, we have, h i
R(ℓaug ) ≤ R(ℓ) + max LLip EGem ∼ Aλ Genm − Gn .
n n
7.3 G R AT I N : G A U S S I A N M I X T U R E M O D E L F O R G R A P H D ATA A U G M E N TAT I O N 85

Theorem 7.3.1 relies on the assumption that the loss function is Lipschitz continuous. This assumption
is realistic, given that the input node features and graph structures in real-world datasets are typically
bounded, i.e., node features are typically normalized or constrained within a fixed range, while graph
structures, represented by adjacency matrices or their normalized forms, have bounded spectral properties,
ensuring a constrained input space. Additionally, we can ensure that the loss function is bounded within
[0, 1] by composing any standard classification loss with a strictly increasing function that maps values to
the interval [0, 1]. We provide the proof of this theorem in Appendix e.1.
A direct implication of Theorem 7.3.1 is that if we chose the right data augmentation strategy h Aλ thati
minimizes the expected distance between original graphs and augmented ones EG∼D ,G∼ e Aλ Ge − G ,
we can guarantee with a high probability that the data augmentation decreases both the Rademacher
complexity and the generalization risk. Specifically, we need enough diversity to ensure a non-zero
difference in the Rademacher complexity compared to the non-augmented case while controlling the
expected distance so that, with high probability, we reduce both the Rademacher complexity and the
overall generalization error η. On the other hand, if the distance is large, we cannot guarantee that data
augmentation will outperform the normal training setting.
The findings of Theorem 7.3.1 hold for all norms defined on the graph input space. Specifically, let
us consider the graph structure space (A, ∥·∥A ) and the feature space (X, ∥·∥X ), where ∥·∥A and ∥·∥X
denote the norms applied to the graph structure and features, respectively. Assuming a maximum number
of nodes per graph, which is a realistic assumption for real-world data, the product space A × X is a
finite-dimensional real vector space, and all the norms are equivalent. Thus, the choice of norm does not
affect the theorem, as long as the Lipschitz constant is adjusted accordingly. Additional details on graph
distance metrics and the comparison between the Lipschitz constants of GCN and GIN are provided in
Appendices e.9 and e.4.
The variety of distance metrics on the original graph space offers different upper bounds, which leads
to distinct criteria for data augmentation based on these metrics to control the upperbound in Theorem
7.3.1. Instead, we propose to shift the focus to the hidden representation space of graphs, where we aim to
derive more consistent and meaningful augmentations. If LLip is taken to be the Lipschitz constant of the
post-readout function only, then Theorem 7.3.1 still applies for the norm on the graph-level embeddings
produced by the readout function, i.e., ∥hGe − hG ∥.
Working at the level of hidden representations of graphs, rather than directly in the graph input space,
offers additional advantages. Hidden representations capture both the structural information and node
features of each graph, enabling augmentation that enhances the generalization of both aspects simul-
taneously. Moreover, node alignment is needed to compare the original and augmented graph, which is
computationally expensive. By operating on hidden representations instead, node alignment becomes
unnecessary. Furthermore, as we will discuss in Section 7.3.4, the effectiveness of augmented data
depends on the specific GNN architecture. By leveraging graph representations learned through a GNN,
we ensure that the augmentation process remains architecture-specific, aligning with the inductive biases
of the chosen model.

7.3.2 Proposed Approach

Based on the theoretical findings,


h it is icrucial to employ a data augmentation technique that effectively
controls the term EG∼D ,G∼
e Aλ ∥ hGe − hG ∥ which measures the expected deviation in the hidden represen-
86 I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

tations of graphs under augmentation, to achieve stronger generalization guarantees. To better understand
this term, we express it in terms of graph representations rather than graphs themselves,
h i h i
EG∼D ,G∼
e Aλ ∥ hGe − hG ∥ = Eh∼δ ,h e ∼Q ∥h − h∥ ,
e
D λ

where Qλ represents the augmentation strategy at the level of hidden representations, replacing Aλ , which
operates at the graph level, and δD : h 7→ N1 ∑nN=1 δhGn (h) represents the Dirac distribution over the training
graph representations, capturing the empirical distribution of graph embeddings. In Proposition 7.3.2, we
upper bound this expected perturbation.

Proposition 7.3.2. Let δD denote the discrete distribution of the training graph representations. Suppose
we sample new augmented graph representations from a distribution Qλ defined on Rd . Then, the following
inequality holds,
h i √ q √ 
Eh∼δ ,he ∼Q ∥h − h∥ ≤ 2 · sup ∥h − h∥
e e KL(δD ∥ Qλ ) + 2 , (7.5)
D λ
h∼δD
e ∼ Qλ
h

where KL(· ∥ ·) denotes the Kullback-Leibler divergence.

We provide a proof of Proposition 7.3.2 in Appendix e.2. A way to control the left side of this in-
equality is to choose a generator Qλ that minimizes both the KL(δD ∥ Qλ ) and the supremum distance
suph∼δ ,he ∼Q ∥h − he ∥. Various universal approximators can be used to minimize KL(δD ∥ Qλ ), including
D λ
generative models like Generative Adversarial Networks (GANs) Yang, Lai, and Lin, 2012. These models
are capable of approximating any probability distribution, making them powerful tools for learning complex
augmentations. However, we specifically choose Gaussian Mixture Models (GMMs), which are well-suited
for this purpose, and can effectively approximate any data distribution, c.f. Theorem 7.3.3. GMMs are
computationally fast compared to other generative approaches, making them suitable for large-scale
graph datasets. As shown in Appendix e.6, GMM-based augmentation yields better results compared
to alternative generative strategies. Moreover, due to the exponential decay of Gaussian distributions,
the supremum distance suph∼δ ,he ∼Q ∥h − h e ∥ is naturally constrained, ensuring a better control of the
hD λ i

expected distance Eh∼δ ,he ∼Q ∥h − h e∥ .


D λ

Theorem 7.3.3. (Goodfellow, Bengio, and Courville, 2016, Page 65) A Gaussian mixture model is a
universal approximator of densities, in the sense that any smooth density can be approximated with any
specific nonzero amount of error by a Gaussian mixture model with enough components.

To achieve this, we first train a standard GNN on the graph classification task using the training
set. Next, we obtain embeddings for all training graphs using the READOUT output, resulting in H =
{hGn s.t. Gn ∈ Dtrain }. These embeddings are used as the basis for generating augmented training
graphs. We then partition the training set Dtrain by classes, such that Dtrain = c Dc where Dc =
S

{Gn ∈ Dtrain , yn = c}. The objective is to learn new graph representations from these embeddings,
and create augmented data for improved training.
We use the EM algorithm to learn the best-fitting GMM for the embeddings of each partition Dc , denoted
as Hc = {hGn s.t. Gn ∈ Dc }. The EM algorithm finds maximum likelihood estimates for each cluster Hc ,
following the procedure described in Bishop and Nasrabadi, 2006.
Once a GMM distribution pc is fitted for each partition Dc , we use this GMM to generate new augmented
data by sampling hidden representations from pc . Each new sample drawn from pc is then assigned the
corresponding partition label c, ensuring that the augmented data inherits the label structure from the
original partitions. After merging the hidden representations of both the original training data and the
7.3 G R AT I N : G A U S S I A N M I X T U R E M O D E L F O R G R A P H D ATA A U G M E N TAT I O N 87

Graph Representations Nodes


Edges

Trained Sampled Augmented Node Attr.

Representations Representations AGGREGATE


COMBINE
READOUT
Output of
REDOUT
Post-Readout

Sample
Fit
Training of Message Passing Layers
Set of train graphs
from the same class

Figure 7.1 – Illustration of GRATIN. Step 1. We first train the GNN on the graph classification task using the training
graphs. Step 2. Next, we utilize the weights from the message passing layers to generate graph
representations for the training graphs. Step 3. A GMM is then fit to these graph representations,
from which we sample new graph representations. Step 4. Finally, we fine-tune the post-readout
function for the graph classification task, using both the original training graphs and the augmented
graph representations. For inference on the test set, we use the message passing weights trained in
Step 1 and the post-readout function weights trained in Step 4.

augmented graph data, we finetune the post-readout function, i.e., the final part of the GNN, which occurs
after the readout function, on the graph classification task. Since the post-readout function consists of a
linear layer followed by a Softmax function, the finetuning process is relatively fast. To evaluate our model
during inference on test graphs, we input the test graphs into the GNN layers trained in the initial step to
compute the hidden graph representations. For the post-readout function, we use the weights obtained
from the second stage of training. Algorithm 4 and Figure 7.1 provide a summary of the GRATIN model.

7.3.3 Time Complexity

One advantage of GRATIN is its efficiency, as it generates new augmented graph representations
with minimal computational time. Unlike baseline methods, which apply augmentation strategies to each
individual training graph (or pair of graphs in Mixup-based approaches) separately, our method learns
the distribution of graph representations across the entire training dataset simultaneously using the EM
algorithm (Ng, 2000). If N = |Dtrain | is the number of training graphs in the dataset, d is the dimension of
graph hidden representations {hG , G ∈ Dtrain }, and K is the number of Gaussian components in the GMM,
then the complexity to fit a GMM on T iterations is O( N · K · T · d2 ) (Yang, Lai, and Lin, 2012). We compare
the data augmentation times of our approach and the baselines in Table e.4. Due to our different training
scheme, i.e., where we first train the message passing layers and then train the post-readout function
after learning the GMM distribution, we have measured the total backpropagation time and compared
it with the backpropagation time of the baseline methods. The training time of baseline models varies
depending on the augmentation strategy used, specifically whether it involves pairs of graphs or individual
graphs. Even in cases where a graph augmentation has a low computational cost for some baselines,
training can still be time-consuming as multiple augmented graphs are required to achieve satisfactory test
accuracy. In contrast, GRATIN generates only one augmented graph per training graph, demonstrating
effective generalization on the test set. Overall, our data augmentation approach is highly efficient during
the sampling of augmented data, with minimal impact on the training time. A complete analysis of the time
complexity of GRATINand the baselines can be found in Appendix e.7.
88 I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

Algorithm 4: Graph classification with GRATIN


1 Inputs: GNN of T layers f (· , θ ) = Ψ ◦ READOUT ◦ g, where g is the composition of message passing layers,

i.e, g = ∪tT=0 {AGGREGATE(t) ◦ COMBINE(t) (·)} and Ψ is the post-readout function, graph classification dataset
D , loss function L;
2 Steps:
3 1. Train GNN f on the training set Dtrain ;
4 2. Use the trained message passing layers and the readout function to generate graph representations
H = {hGn s.t. Gn ∈ Dtrain } for the training set;
3. Partition the training set Dtrain by classes, such that Dtrain = c Dc where Dc = {Gn ∈ Dtrain , yn = c};
S
5
6 foreach c ∈ {0, . . . , C } do
7 3.1. Fit a GMM distribution pc on the graph representations Hc = {hGn s.t. Gn ∈ Dc };
8 3.2. Sample new graph representation H fc = {he s.t. he ∼ pc } from the distribution pc ;
9 3.3. Include the sampled representations H fc with trained representations Hc = Hc ∪ H fc ;
10 4. Finetune the post-readout function Ψ on the graph classification task directly on the new training set
H = ∪c Hc ;

7.3.4 Analyzing the Generalization Ability of the Augmented Graphs via Influence Functions

We use influence functions (Koh and Liang, 2017; Kong, Shen, and Huang, 2021; Law, 1986) to
understand the impact of augmented data on the model performance on the test set and thus motivate
the use of a data augmentation strategy, which is specific to the model architecture and model weights. In
Theorem 7.3.4, we derive a closed-form formula for the impact of adding an augmented graph Genm on the
GNN’s performance on a test graph Gktest , where the GNN is trained solely on the training set, without the
augmented graph.

Theorem 7.3.4. Given a test graph Gk , let θ̂ = arg minθ L be the GNN parameters that minimize the
aug
objective function in (7.1). The impact of upweighting the objective function L to Ln,m = L + ϵn,m ℓ(Genm , θ ),
where Genm is an augmented graph candidate of the training graph Gn and ϵn,m is a sufficiently small
perturbation parameter, on the model performance on the test graph Gktest is given by

dℓ(Gktest , θ̂ϵn,m )
= −∇θ ℓ(Gktest , θ̂ )H− 1
∇θ ℓ(Genm , θ̂ ), (7.6)
dϵn,m θ̂

aug aug
where θ̂ϵn,m = arg minθ Ln,m denotes the parameters that minimize the upweighted objective function Ln,m
and Hθ̂ = ∇2θ L(θ̂ ) is the Hessian matrix of the loss w.r.t. the model parameters.

We provide the proof of Theorem 7.3.4 in Appendix e.3. The influence scores are useful for evaluating
the effectiveness of the augmented data on each test graph. The strength of influence function theory lies
in its ability to analyze the effect of adding augmented data to the training set without actually retraining
on this data. As noticed, these influence scores depend not only on the augmented graphs themselves
but also on the model’s weights and architecture. This highlights the need for a graph data augmentation
strategy tailored specifically to the GNN backbone in use, as opposed to traditional techniques like
DropNode, DropEdge, and G -Mixup, which are general-purpose methods that can be applied with any
GNN architecture.
Theorem 7.3.4 is valid for any differentiable loss function. More specifically, if the chosen loss is the
cross entropy or the negative log-likelihood, then the hessian matrix corresponds to the Fisher information
matrix Barshan, Brunet, and Dziugaite, 2020; Lee, Kim, and Ro, 2022. Consequently, the norm of H− θ̂
1
,
i.e., the inverse of the hessian matrix, can be bounded above using the Cramér–Rao inequality Nielsen,
7.4 E X P E R I M E N TA L R E S U LT S 89

MUTAG PROTEINS DD
40
GCN 15 GCN 250 GIN
30 GIN GIN 200 GCN
Density

Density

Density
10 150
20
100
10 5
50
0 0 0
0.02 0.00 0.02 0.04 0.06 0.08 0.0 0.1 0.2 0.3 0.1 0.0 0.1 0.2 0.3
Influence Scores Influence Scores Influence Scores
Figure 7.2 – The density of the average influence scores of each augmented data on the test set.

2013. Therefore, a trivial case where the norm of influence scores is zero arises when the gradient of
the loss function with respect to the input graphs vanishes. This scenario, for instance, can occur in the
DD dataset when using GIN. A detailed analysis of this phenomenon is provided in Section 7.4. In these
cases, data augmentation becomes ineffective, having minimal impact on the GNN’s ability to generalize.
We can measure the average influence I(Genm ) of a augmented graph Genm on the test set by averaging the
derivatives as follows,
−1 dℓ(Gktest , θ̂ϵn,m )
|Dtest | G test∑
I(Genm ) = .
∈D
dϵn,m
k test

A negative value of I(Genm ) indicates that adding the augmented data to the training set would increase the
prediction loss on the test set, negatively affecting the GNN’s generalization. In contrast, a good augmented
graph is one with a positive I(Genm ), indicating improved generalization. In Figure 7.2, we present the
density of the average influence scores of each augmented data on the test set.

7.3.5 Fisher-Guided GMM Augmentation

Using influence scores, we can further improve the generalization of the GNN by filtering candidate
augmented representations. The process consists of three key stages. (i) Primary GNN training: The
GNN model is first trained on the original training set without incorporating any augmented graphs. (ii)
Augmentation and filtering: A pool of candidate augmented graph representations is generated using a
data augmentation strategy based on GMMs. The influence of each augmented representation on the
validation set is then computed using influence scores, enabling us to rank the augmented graphs by
their impact. (iii) Filtering: Finally, we combine a subset of the highest-ranked augmented graphs with the
original training set to train the post-readout function. Our experiments in Appendix e.12 demonstrate that
this two-phase training paradigm improves generalization across various datasets and GNN architectures.

7.4 E X P E R I M E N TA L R E S U LT S

In this section, we present our results and analysis. Our experimental setup is described in Appendix
e.11.
On the Generalization of GNNs. In Tables 7.1 and 7.2, we compare the test accuracy of our data
augmentation strategy against baseline methods. We trained all baseline models using the same train/vali-
dation/test splits, GNN architectures, and hyperparameters to ensure a fair comparison. It is worth noting
that the baselines exhibit high standard deviations, which is a common characteristic in graph classification
90 I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

Table 7.1 – Classification accuracy (± std) on different benchmark graph classification datasets for the data aug-
mentation baselines based on the GCN backbone. The higher the accuracy (in %) the better the model.
Highlighted are the first, second best results.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

No Aug. 73.00±4.94 47.73±2.64 73.92±5.09 69.99±5.35 69.69±2.89


DropEdge 71.70±5.42 45.67±2.46 73.39±8.86 70.07±3.86 69.35±3.37
DropNode 74.00±3.44 43.80±3.54 73.89±8.53 69.81±4.61 69.01±3.95
SubMix 72.70±5.59 46.00±2.44 77.13±9.69 67.57±4.56 70.11±4.48
G -Mixup 72.10±3.27 48.33±3.06 88.77±5.71 65.68±5.03 61.20±3.88
GeoMix 69.69±3.37 49.80±4.71 74.39±7.37 69.63±5.37 68.50±3.74
GRATIN 71.00±4.40 49.82±4.26 76.05±6.74 70.97±5.07 71.90±2.81

Table 7.2 – Classification accuracy (± std) on different benchmark graph classification datasets for the data aug-
mentation baselines based on the GIN backbone. The higher the accuracy (in %) the better the model.
Highlighted are the first, second best results.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

No Aug. 70.30±3.66 48.53±4.05 83.42±2.12 69.54±3.61 68.00±3.18


DropEdge 70.40±4.03 46.80±3.91 74.88±9.62 68.27±5.21 67.82±4.46
DropNode 70.30±3.49 45.20±4.24 75.53±7.89 65.40±4.71 69.01±3.95
SubMix 72.50±4.98 48.13±2.12 81.90±9.21 70.44±2.58 68.59±5.04
G -Mixup 70.70±3.10 47.73±4.95 87.77±7.48 68.82±3.48 63.91±2.09
GeoMix 70.60±4.61 47.20±3.75 81.90±7.55 69.80±5.33 68.34±5.30
GRATIN 71.70±4.24 49.20±2.06 88.83±5.02 71.33±5.04 68.61±4.62

tasks. Unlike node classification, graph classification is known to have a larger variance in performance
metrics (Duval and Malliaros, 2022; Errica et al., 2019). Overall, our proposed approach consistently
achieves the best or highly competitive performance for most of the datasets.
Additionally, we observed that the results of the baseline methods vary depending on the GNN backbone,
motivating further investigation using influence functions. As demonstrated in Theorem 7.3.4, the gradient,
and more generally, the model architecture, significantly influence how augmented data impacts the model’s
performance on the test set.
Robustness to Structure Corruption. Besides generalization, we assess the robustness of our data
augmentation strategy, following the methodology outlined by (Zeng et al., 2024). Specifically, we test the
robustness of data augmentation strategies against graph structure corruption by randomly removing or
adding 10% or 20% of the edges in the training set. By corrupting only the training graphs, we introduce a
distributional shift between the training and testing datasets. This approach allows us to evaluate GRATIN’s
ability to generalize well and predict the labels of test graphs, which can be considered OOD examples.
The results of these experiments are presented in Table 7.3 for the IMDB-BIN, IMDB-MUL, PROTEINS,
7.5 C O N C L U S I O N 91

Table 7.3 – Robustness against structure corruption: We present the Classification accuracy (± standard deviation).
We highlighted the best data augmentation strategy bold. For this experiment, we use the GCN backbone.
Noise Budget 10% 20%

Dataset IMDB-BIN IMDB-MUL PROTEINS DD IMDB-BIN IMDB-MUL PROTEINS DD

DropNode 66.40±5.51 44.46±2.13 69.18±4.87 65.79±3.23 64.80±5.01 43.06±2.86 67.73±6.43 64.35±4.56


DropEdge 66.70±5.10 43.80±3.11 69.36±5.90 68.42±4.76 63.20±6.30 41.80±3.15 68.10±5.05 67.06±2.53
SubMix 69.30±3.76 46.73±2.67 69.80±4.73 68.04±7.64 63.70±5.64 43.73±3.60 69.09±4.58 59.18±6.29
GeoMix 72.20±5.19 49.20±4.31 70.25±4.75 68.00±3.64 70.90±3.85 48.86±5.18 68.36±6.01 67.31±3.91
G -Mixup 68.30±5.13 45.53±4.12 61.71±5.81 51.26±8.76 63.20±5.54 44.00±4.63 46.63±5.05 43.71±7.12
NoisyGNN 70.50±4.71 40.66±3.12 69.45±4.32 64.18±5.71 63.50±5.43 38.66±4.12 69.99±3.78 63.24±5.02
GRATIN 72.80±2.99 49.36±4.53 70.61±4.30 68.68±3.72 73.10±3.04 49.53±3.54 70.32±4.04 69.01±3.09

and DD datasets. As noted, our data augmentation strategy exhibits the best test accuracy in all cases and
improves model robustness against structure corruption.
Influence Functions. In Figure 7.2, we show the density distribution of the average influence of
augmented data sampled using GRATIN. These findings are consistent with the empirical results presented
in Tables 7.1. For the MUTAG and PROTEINS datasets, we observe that GRATIN’s data augmentation
has a positive impact on both GCN and GIN models. In contrast, for the DD dataset, GRATIN shows no
effect on GIN, while it generates many augmented samples with positive values of the influence scores
on GCN, thereby enhancing its performance. This behavior is consistent with the baselines, as most
graph data augmentation strategies tend to enhance test accuracy more significantly for GCN than for
GIN when applied to DD. This is an interesting phenomenon and worthy of deeper analysis. For the DD
dataset, when using the GIN model at inference, we observe Softmax saturation, where the predicted class
probabilities approach extreme values (close to 0 or 1), c.f. Appendix e.13. We suspect this saturation to
happen due to the large average number of nodes in DD and the fact that GIN does not normalize the
node representations, which may lead to a graph representation with large norms. This saturation results
in the model making predictions with very high confidence. Consequently, the gradient of the loss function
with respect to the input graphs eventually converges to 0. In such cases, the influence scores become
negligible, as explained in Section 7.3.4.
Configuration Models. As part of an ablation study, we propose a simple yet effective graph augmenta-
tion strategy inspired by configuration
h modelsi (Newman, 2013). As shown in Theorem 7.3.1, the objective
is to control the term EG∼D ,G∼
e Aλ ∥ hGe − hG ∥ , which can be achieved by regulating the distance between
h i
the original and the sampled graph within the input manifold, i.e., EG∼D ,G∼ e Aλ ∥ G − G∥ . The approach
e
involves generating a sampled version of each training graph by randomly breaking existing edges into
half-edges with probability r and then randomly connecting half-edges until all edges are connected. The
strength of this method lies in its simplicity and in preserving the degree distribution. Ifh the distance
i norm is
the L1 distance between adjacency matrices, |E |r is an upper bound of EG∼D ,G∼
2
e Aλ ∥ G − G∥ , where |E |
e
is the average number of edges. The results of this experiment are available in Appendix e.5.

7.5 CONCLUSION

We introduced GRATIN, a novel approach for graph data augmentation that enhances both the general-
ization and robustness of GNNs. Our method uses Gaussian Mixture Models (GMMs) applied at the output
92 I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

level of the Readout function, an approach motivated by theoretical findings. Using the universal approxi-
mation property of GMMs, we can sample new graph representations to effectively control the upper bound
of the Rademacher complexity, ensuring improved generalization of GNNs. Through extensive experiments
on widely used datasets, we demonstrated that our approach not only exhibits strong generalization ability
but also maintains robustness against structural perturbations. An additional advantage of our approach is
its efficiency in terms of time complexity. Unlike baselines that generate augmented data for each individual
or pair of training graphs, GRATIN fits the GMM to the entire training dataset at once, allowing for fast graph
data augmentation without incurring significant additional backpropagation time.
Part V

CONCLUSION
CONCLUSION
8

95
96 CONCLUSION

8.1 S U M M A RY O F C O N T R I B U T I O N S

8.1.1 Representation learning

Chapter 3 introduced Centrality Graph Shift Operators (CGSOs), a novel class of graph shift operators.
These operators incorporate both local and global centrality measures, such as degree, k-core, walk-count,
and PageRank centralities, enabling flexible propagation. Spectral analysis demonstrated that CGSOs more
accurately capture community structure compared to traditional degree-based GSOs, resulting in consistent
performance gains across synthetic clustering tasks and a variety of node classification benchmarks.
In Chapter 4, we presented Adaptive-Depth Message Passing GNNs (ADMP-GNNs), which eliminate the
need to predefine the number of message-passing steps. Instead, ADMP-GNNs employ centrality-informed
policies to dynamically learn the optimal propagation depth for each node. Empirical results highlight that
this node-specific flexibility leads to improvements in predictive accuracy.

8.1.2 Robustness

Chapter 5 introduced a new formulation of expected robustness for graphs, providing a framework
to evaluate GNNs vulnerability to adversarial perturbations. We derived theoretical bounds connecting
a model’s robustness to the orthonormality of its weight matrices, revealing that orthonormality leads
to improved resistance against attacks. Building on this insight, we proposed the Graph Convolutional
Orthonormal Robust Network (GCORN), a variant of the classical GCN that enforces strict orthonormality.
Complementing the training-time defense, Chapter 6 introduced RobustCRF, a post-hoc, model-agnostic
method based on conditional random fields that make predictions more robust during inference. RobustCRF
improves performance on both clean and adversarially perturbed data, and could me combined with other
baselines.

8.1.3 Generalization

Finally, Chapter 7 addresses the generalization of GNNs through the introduction of GRATIN, a novel
graph data augmentation technique. GRATIN fits a Gaussian Mixture Model to the latent representations
produced by a trained GNN, enabling the generation of additional graph representations in that learned
space. The method is theoretically supported via an upper bound on the Rademacher complexity. Empir-
ically, GRATIN improves test accuracy on the graph classification task and demonstrates robustness to
structural noise, all by maintaining low computational overhead during both augmentation and training.

8.2 B R O A D E R I M PA C T

The methods introduced in this thesis aim to enhance the robustness, and generalization of GNNs, with
the goal of making them more reliable and adaptable for real-world use.
In terms of representation learning, the development of Centrality Graph Shift Operators (CGSOs)
provides more flexibility for GNNs to capture structural information in graphs. This has broad implications
for domains such as social network analysis (e.g., Twitter), music recommendation systems, e.g., Deezer
(Salha-Galvan, 2022) and biological network modeling (Liu et al., 2023), where understanding the im-
portance and roles of nodes within a graph is crucial. By aligning message passing with global or local
structural properties, CGSOs enable more interpretable and efficient graph representations.
8.3 F U T U R E D I R E C T I O N S 97

For robustness, techniques like GCORN and RobustCRF help mitigate the vulnerabilities of GNNs to
adversarial manipulations. This is especially significant in safety-critical applications such as navigation
systems, e.g., Google Maps (Lange and Perez, 2020) or fraud detection. These methods contribute to the
broader effort of building trustworthy machine learning systems, independent of specific attack models or
retraining procedures.
On the front of generalization, the GRATIN framework addresses a key challenge in real-world GNN
deployment: the ability to perform well under limited data or when confronted with distributional shifts.
Applications in evolving social networks and drug discovery (Kaur, Kukreja, and Singh, 2024; Liu and Xie,
n.d.), where molecular graphs often vary, stand to benefit from GRATIN’s ability to augment training data.
Importantly, this is achieved without introducing significant computational cost.
Together, the contributions of this thesis open up new directions toward GNNs that are not only more
accurate but also more robust, interpretable, and deployable in real-world scenarios. At the same time, they
highlight new ethical considerations, such as the potential for structural biases, or the risk of overfitting to
learned latent representations. These call for attention to fairness in future GNN research and deployment.

8.3 FUTURE DIRECTIONS

This thesis opens several avenues for further investigation. Below we highlight a few especially promising
directions:
— Generalization to Node Classification under Non-IID Settings. Most classical generalization
bounds, such as those based on Rademacher complexity or McDiarmid’s inequality, assume that
training and test examples are drawn independently and identically distributed. While this assumption
is reasonable for graph-level tasks, it breaks down in semi-supervised node classification, i.e., nodes
are connected by edges, and label information propagates through the same graph on which we test.
Developing a theory of generalization for node classification task should accounts for this inherent
dependency structure in order to derive new concentration inequalities. Such a framework would
enable rigorous bounds on node-level error and guide the design of augmentation or regularization
techniques for non-IID node distributions.
— 3D Geometric Deep Learning. An exciting direction is to develop E(3) and SE(3) Equivariant
Graph Shift Operators tailored to three-dimensional data Duval et al., 2023a. Current GNNs achieve
permutation equivariance, but modeling 3D structures (e.g., molecular conformations, point clouds,
meshes) demands invariance to rotations and translations. By designing GSOs that explicitly
incorporate spatial coordinates, edge angles, and embedding them into the message-passing scheme
of Equivariant GNN layers—we can respect the symmetries of the Special Euclidean group and
enhance the network’s ability to capture geometric relations. Such 3D-specific GSOs could update
both node features and coordinate embeddings at each layer, enabling more accurate property
prediction in physics and chemistry, and more faithful modeling of complex engineering or biological
structures.
— Transformers and Large Language Models for Structured Data. Recent advances in Large
Language Models (LLMs) open the door to in-context learning for discrete, structured data tasks. By
prompting an LLM with examples of graph transformations or denoising steps, one could perform
tasks like diffusion-based graph generation or adaptive message passing without extensive retraining.
Exploring how to leverage LLMs for link prediction, node classification, or discrete diffusion on graphs
represents a novel, training-free paradigm for graph machine learning.
BIBLIOGRAPHY

Abbahaddou, Yassine, Sofiane Ennadir, Johannes F. Lutzeyer, Michalis Vazirgiannis, and Henrik Boström
(2024a). « Bounding the Expected Robustness of Graph Neural Networks Subject to Node Feature
Attacks. » In: The Twelfth International Conference on Learning Representations (cit. on pp. 5, 68, 69,
74–76, 81, 186).
Abbahaddou, Yassine, Sofiane Ennadir, Johannes F. Lutzeyer, Michalis Vazirgiannis, and Fragkiskos D.
Malliaros (2024b). « Rethinking Robustness in Graph Neural Networks: A Post-Hoc Approach With
Conditional Random Fields. » In: arXiv preprint arXiv:2411.05399 (cit. on p. 5).
Abbahaddou, Yassine, Johannes Lutzeyer, and Michalis Vazirgiannis (2023). « Graph Neural Networks on
Discriminative Graphs of Words. » In: NeurIPS 2023 Workshop: New Frontiers in Graph Learning (cit. on
p. 9).
Abbahaddou, Yassine, Fragkiskos D. Malliaros, Johannes F. Lutzeyer, Amine Mohamed Aboussalah, and
Michalis Vazirgiannis (2025). « Graph Neural Network Generalization With Gaussian Mixture Model
Based Augmentation. » In: The Forty-Second International Conference on Machine Learning (cit. on
p. 5).
Abbahaddou, Yassine, Fragkiskos D. Malliaros, Michalis Vazirgiannis, and Johannes F. Lutzeyer (2024c).
« Centrality Graph Shift Operators for Graph Neural Networks. » In: arXiv preprint arXiv:2411.04655
(cit. on p. 4).
Abbe, Emmanuel (2018). « Community detection and stochastic block models: recent developments. » In:
Journal of Machine Learning Research 18.177, pp. 1–86 (cit. on p. 35).
Ablin, Pierre and Gabriel Peyré (2022). « Fast and accurate optimization on the orthogonal manifold without
retraction. » In: International Conference on Artificial Intelligence and Statistics. PMLR, pp. 5636–5657
(cit. on p. 60).
Aboussalah, Amine Mohamed, Minjae Kwon, Raj G Patel, Cheng Chi, and Chi-Guhn Lee (2023). « Re-
cursive Time Series Data Augmentation. » In: The Eleventh International Conference on Learning
Representations (cit. on p. 82).
Adamic, Lada A and Natalie Glance (2005). « The political blogosphere and the 2004 US election: divided
they blog. » In: Proceedings of the 3rd international workshop on Link discovery, pp. 36–43 (cit. on
p. 170).
Albert, Reka and Albert-Laszlo Barabasi (2002). « Statistical mechanics of complex networks. » In: Reviews
of Modern Physics (cit. on p. 35).
Alchihabi, Abdullah, Qing En, and Yuhong Guo (2023). « Efficient Low-Rank GNN Defense Against
Structural Attacks. » In: 2023 IEEE International Conference on Knowledge Graph (ICKG). IEEE, pp. 1–8
(cit. on p. 68).
Anderson, Brendon G and Somayeh Sojoudi (2022). « Certified robustness via locally biased randomized
smoothing. » In: Learning for Dynamics and Control Conference. PMLR, pp. 207–220 (cit. on p. 69).
Andrews, Mark and Thom Baguley (2017). « Bayesian data analysis. » In: The Cambridge encyclopedia of
child development, pp. 165–169 (cit. on p. 72).
Anil, Cem, James Lucas, and Roger Grosse (2019). « Sorting out Lipschitz function approximation. » In:
International Conference on Machine Learning. PMLR, pp. 291–301 (cit. on p. 60).
Backstrom, Lars and Jure Leskovec (2011). « Supervised random walks: predicting and recommending
links in social networks. » In: Proceedings of the fourth ACM international conference on Web search
and data mining, pp. 635–644 (cit. on p. 3).

99
100 BIBLIOGRAPHY

Barabási, Albert-László and Réka Albert (1999). « Emergence of scaling in random networks. » In: science
286.5439, pp. 509–512 (cit. on p. 35).
Barshan, Elnaz, Marc-Etienne Brunet, and Gintare Karolina Dziugaite (2020). « Relatif: Identifying explana-
tory training samples via relative influence. » In: International Conference on Artificial Intelligence and
Statistics. PMLR, pp. 1899–1909 (cit. on p. 88).
Benson, Austin R, David F Gleich, and Jure Leskovec (2016). « Higher-order organization of complex
networks. » In: Science 353.6295, pp. 163–166 (cit. on pp. 16, 32, 48).
Bernoulli, Jakob (1713). Ars conjectandi: opus posthumum: accedit Tractatus de seriebus infinitis; et
Epistola gallice scripta de ludo pilae reticularis. Impensis Thurnisiorum (cit. on p. 61).
Bevilacqua, Beatrice, Fabrizio Frasca, Derek Lim, Balasubramaniam Srinivasan, Chen Cai, Gopinath
Balamurugan, Michael M. Bronstein, and Haggai Maron (2022). « Equivariant Subgraph Aggregation
Networks. » In: International Conference on Learning Representations. URL: [Link]
forum?id=dFbKQaRk15w (cit. on p. 20).
Bishop, Christopher M and Nasser M Nasrabadi (2006). Pattern recognition and machine learning. Vol. 4.
Springer (cit. on p. 86).
Björck, Å. and C. Bowie (1971). « An Iterative Algorithm for Computing the Best Estimate of an Orthogonal
Matrix. » In: SIAM Journal on Numerical Analysis, pp. 358–364 (cit. on pp. 60, 155).
Blei, David M, Alp Kucukelbir, and Jon D McAuliffe (2017). « Variational inference: A review for statisticians. »
In: Journal of the American statistical Association 112.518, pp. 859–877 (cit. on p. 72).
Blondel, Vincent D, Jean-Loup Guillaume, Renaud Lambiotte, and Etienne Lefebvre (2008). « Fast unfolding
of communities in large networks. » In: Journal of statistical mechanics: theory and experiment 2008.10,
P10008 (cit. on p. 37).
Bojchevski, Aleksandar, Johannes Gasteiger, and Stephan Günnemann (2020). « Efficient robustness
certificates for discrete data: Sparsity-aware randomized smoothing for graphs, images and more. » In:
International Conference on Machine Learning. PMLR, pp. 1003–1013 (cit. on pp. 57, 61, 64).
Bojchevski, Aleksandar, Johannes Gasteiger, Bryan Perozzi, Amol Kapoor, Martin Blais, Benedek Rózem-
berczki, Michal Lukasik, and Stephan Günnemann (2020). « Scaling graph neural networks with approxi-
mate pagerank. » In: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge
Discovery & Data Mining, pp. 2464–2473 (cit. on p. 31).
Bojchevski, Aleksandar and Stephan Günnemann (2019a). « Certifiable robustness to graph perturba-
tions. » In: Advances in Neural Information Processing Systems 32 (cit. on p. 19).
– (2019b). « Certifiable robustness to graph perturbations. » In: Advances in Neural Information Processing
Systems 32 (cit. on p. 57).
Bojchevski, Aleksandar, Johannes Klicpera, and Stephan Günnemann (2020). « Efficient Robustness
Certificates for Discrete Data: Sparsity-Aware Randomized Smoothing for Graphs, Images and More. »
In: CoRR abs/2008.12952. arXiv: 2008.12952. URL: [Link] (cit. on p. 69).
Bolukbasi, Tolga, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama (2017). « Adaptive neural networks
for efficient inference. » In: International Conference on Machine Learning. PMLR, pp. 527–536 (cit. on
p. 41).
Borgatti, Stephen P (2005). « Centrality and network flow. » In: Social networks 27.1, pp. 55–71 (cit. on
p. 14).
Bornholdt, Stefan and Heinz Georg Schuster (2001). « Handbook of graphs and networks. » In: From
Genome to the Internet, Willey-VCH (2003 Weinheim) (cit. on p. 41).
Brin, Sergey and Lawrence Page (1998). « The anatomy of a large-scale hypertextual web search engine. »
In: Computer networks and ISDN systems 30.1-7, pp. 107–117 (cit. on pp. 14, 15, 29, 32, 48).
BIBLIOGRAPHY 101

Brody, Shaked, Uri Alon, and Eran Yahav (2022). « How Attentive are Graph Attention Networks? » In:
International Conference on Learning Representations. URL: [Link]
F72ximsx7C1 (cit. on pp. 17, 31, 33, 38).
Brouwer, Andries E and Willem H Haemers (2011). Spectra of graphs. Springer Science & Business Media
(cit. on p. 14).
Bruna, Joan, Wojciech Zaremba, Arthur Szlam, and Yann LeCun (2014). « Spectral networks and locally
connected networks on graphs. » In: International Conference on Learning Representations (ICLR)
(cit. on p. 31).
Buffelli, Davide, Pietro Liò, and Fabio Vandin (2022). « Sizeshiftreg: a regularization method for improving
size-generalization in graph neural networks. » In: Advances in Neural Information Processing Systems
35, pp. 31871–31885 (cit. on p. 81).
Buser, Peter (1982). « A note on the isoperimetric constant. » In: Annales scientifiques de l’École normale
supérieure. Vol. 15. 2, pp. 213–230 (cit. on p. 13).
Cai, Xing Shi, Pietro Caputo, Guillem Perarnau, and Matteo Quattropani (2023). « Rankings in directed
configuration models with heavy tailed in-degrees. » In: The Annals of Applied Probability 33.6B,
pp. 5613–5667 (cit. on p. 33).
Cao, Defu, Yujing Wang, Juanyong Duan, Ce Zhang, Xia Zhu, Congrui Huang, Yunhai Tong, Bixiong Xu,
Jing Bai, Jie Tong, et al. (2020). « Spectral temporal graph neural network for multivariate time-series
forecasting. » In: Advances in neural information processing systems 33, pp. 17766–17778 (cit. on p. 41).
Carlini, Nicholas and David A. Wagner (2016). « Towards Evaluating the Robustness of Neural Networks. »
In: 2017 IEEE Symposium on Security and Privacy (SP), pp. 39–57 (cit. on p. 71).
Carmon, Yair, Aditi Raghunathan, Ludwig Schmidt, Percy Liang, and John C. Duchi (2019). « Unlabeled
Data Improves Adversarial Robustness. » In: ArXiv abs/1905.13736 (cit. on p. 69).
Castro-Correa, Jhon A., Jhony H. Giraldo, Mohsen Badiey, and Fragkiskos D. Malliaros (2024a). « Gegen-
bauer Graph Neural Networks for Time-Varying Signal Reconstruction. » In: IEEE Transactions on Neural
Networks and Learning Systems 35.9, pp. 11734–11745 (cit. on p. 41).
– (2024b). « Gegenbauer Graph Neural Networks for Time-Varying Signal Reconstruction. » In: IEEE
Transactions on Neural Networks and Learning Systems 35.9, pp. 11734–11745 (cit. on p. 81).
Cheeger, Jeff (1970b). « A lower bound for the smallest eigenvalue of the Laplacian. » In: Problems in
Analysis. Princeton University Press, pp. 195–200 (cit. on p. 14).
– (1970a). « A lower bound for the smallest eigenvalue of the Laplacian. » In: Problems in analysis
625.195-199, p. 110 (cit. on p. 13).
Chen, Chaofan, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su (2019). « This
looks like that: deep learning for interpretable image recognition. » In: Advances in neural information
processing systems 32 (cit. on p. 67).
Chen, Liqun, Zhe Gan, Yu Cheng, Linjie Li, Lawrence Carin, and Jingjing Liu (2020a). « Graph optimal
transport for cross-domain alignment. » In: International Conference on Machine Learning. PMLR,
pp. 1542–1553 (cit. on p. 21).
Chen, Ming, Zhewei Wei, Zengfeng Huang, Bolin Ding, and Yaliang Li (2020b). « Simple and deep graph
convolutional networks. » In: International conference on machine learning. PMLR, pp. 1725–1735 (cit. on
pp. 43, 50).
Chen, Yongqiang, Han Yang, Yonggang Zhang, Kaili Ma, Tongliang Liu, Bo Han, and James Cheng (2022).
« Understanding and improving graph injection attack by promoting unnoticeability. » In: International
Conference of Learning Representations (cit. on p. 56).
Cheng, Wuxinlin, Chenhui Deng, Zhiqiang Zhao, Yaohui Cai, Zhiru Zhang, and Zhuo Feng (2021). « Spade:
A spectral method for black-box adversarial robustness evaluation. » In: International Conference on
Machine Learning. PMLR, pp. 1814–1824 (cit. on pp. 61, 71).
102 BIBLIOGRAPHY

Chi, Cheng, Amine Mohamed Aboussalah, Elias B. Khalil, Juyoung Wang, and Zoha Sherkat-Masoumi
(2022). « A deep reinforcement learning framework for column generation. » In: Proceedings of the 36th
International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA
(cit. on p. 81).
Chien, Eli, Jianhao Peng, Pan Li, and Olgica Milenkovic (2020). « Adaptive universal generalized pagerank
graph neural network. » In: arXiv preprint arXiv:2006.07988 (cit. on pp. 10, 43, 50).
Chung, F. R. K. (1997). Spectral Graph Theory. American Mathematical Society (cit. on pp. 29, 33).
Cisse, Moustapha, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier (2017). « Parseval
networks: Improving robustness to adversarial examples. » In: International Conference on Machine
Learning. PMLR, pp. 854–863 (cit. on pp. 60, 63, 64).
Clauset, Aaron, Mark EJ Newman, and Cristopher Moore (2004). « Finding community structure in very
large networks. » In: Physical review E 70.6, p. 066111 (cit. on p. 37).
Clifford, Peter (1990). « Markov random fields in statistics. » In: Disorder in physical systems: A volume in
honour of John M. Hammersley, pp. 19–32 (cit. on p. 69).
Cormen, Thomas H, Charles E Leiserson, Ronald L Rivest, and Clifford Stein (2022). Introduction to
algorithms. MIT press (cit. on p. 34).
Corso, Gabriele, Luca Cavalleri, Dominique Beaini, Pietro Liò, and Petar Veličković (2020). « Principal
neighbourhood aggregation for graph nets. » In: Advances in Neural Information Processing Systems 33,
pp. 13260–13271 (cit. on p. 38).
Corso, Gabriele, Hannes Stärk, Bowen Jing, Regina Barzilay, and Tommi Jaakkola (2022). « Diffdock:
Diffusion steps, twists, and turns for molecular docking. » In: arXiv preprint arXiv:2210.01776 (cit. on
pp. 41, 81).
Cvetković, Dragoš, Peter Rowlinson, and Slobodan Simić (2009). An Introduction to the Theory of Graph
Spectra. London Mathematical Society Student Texts. Cambridge University Press. DOI: 10 . 1017 /
CBO9780511801518 (cit. on p. 11).
Cvetkovic, Dragos M, Michael Doob, and Horst Sachs (1980). « Spectra of graphs. Theory and application. »
In: (cit. on p. 11).
Cvetković, Dragoš and Slobodan K Simić (2009). « Towards a spectral theory of graphs based on the
signless Laplacian, I. » In: Publications de l’Institut Mathematique 85.105, pp. 19–33 (cit. on p. 12).
– (2010). « Towards a spectral theory of graphs based on the signless Laplacian, II. » In: Linear Algebra
and its Applications 432.9, pp. 2257–2272 (cit. on p. 38).
Dabouei, Ali, Sobhan Soleymani, Fariborz Taherkhani, and Nasser M Nasrabadi (2021). « Supermix:
Supervising the mixing data augmentation. » In: Proceedings of the IEEE/CVF conference on computer
vision and pattern recognition, pp. 13794–13803 (cit. on p. 82).
Dai, Hanjun, Hui Li, Tian Tian, Xin Huang, Lin Wang, Jun Zhu, and Le Song (2018). « Adversarial Attack
on Graph Structured Data. » In: Proceedings of the 35th International Conference on Machine Learning,
pp. 1115–1124 (cit. on pp. 55, 56, 67, 68).
Dasoulas, George, Johannes F. Lutzeyer, and Michalis Vazirgiannis (2021b). « Learning Parametrised
Graph Shift Operators. » In: International Conference on Learning Representations (cit. on pp. 34, 121).
– (2021a). « Learning Parametrised Graph Shift Operators. » In: ArXiv abs/2101.10050. URL: https :
//[Link]/CorpusID:231698861 (cit. on p. 12).
Defferrard, Michaël, Xavier Bresson, and Pierre Vandergheynst (2016). « Convolutional neural networks on
graphs with fast localized spectral filtering. » In: Advances in neural information processing systems 29
(cit. on p. 31).
Dempster, Arthur P, Nan M Laird, and Donald B Rubin (1977). « Maximum likelihood from incomplete
data via the EM algorithm. » In: Journal of the royal statistical society: series B (methodological) 39.1,
pp. 1–22 (cit. on p. 186).
BIBLIOGRAPHY 103

Deng, Chenhui, Xiuyu Li, Zhuo Feng, and Zhiru Zhang (2022). « Garnet: Reduced-rank topology learning
for robust and scalable graph neural networks. » In: Learning on Graphs Conference. PMLR, pp. 3–1
(cit. on p. 68).
Dey, Palash, Suman Kalyan Maity, Sourav Medya, and Arlei Silva (2020). « Network Robustness via Global
k-cores. » In: arXiv preprint arXiv:2012.10036 (cit. on p. 15).
Donath, William E and Alan J Hoffman (1973). « Lower bounds for the partitioning of graphs. » In: IBM
Journal of Research and Development 17.5, pp. 420–425 (cit. on p. 29).
Du, Simon S, Kangcheng Hou, Russ R Salakhutdinov, Barnabas Poczos, Ruosong Wang, and Keyulu Xu
(2019). « Graph neural tangent kernel: Fusing graph neural networks with graph kernels. » In: Advances
in neural information processing systems 32 (cit. on pp. 19, 81).
Duchi, John C, Peter L Bartlett, and Martin J Wainwright (2012). « Randomized smoothing for stochastic
optimization. » In: SIAM Journal on Optimization 22.2, pp. 674–701 (cit. on p. 69).
Duval, Alexandre and Fragkiskos Malliaros (2022). « Higher-order Clustering and Pooling for Graph Neural
Networks. » In: Proceedings of the 31st ACM International Conference on Information & Knowledge
Management, 426–435 (cit. on p. 90).
Duval, Alexandre, Simon V Mathis, Chaitanya K Joshi, Victor Schmidt, Santiago Miret, Fragkiskos D
Malliaros, Taco Cohen, Pietro Liò, Yoshua Bengio, and Michael Bronstein (2023a). « A hitchhiker’s guide
to geometric gnns for 3d atomic systems. » In: arXiv preprint arXiv:2312.07511 (cit. on p. 97).
Duval, Alexandre, Victor Schmidt, Alex Hernández-García, Santiago Miret, Fragkiskos D. Malliaros, Yoshua
Bengio, and David Rolnick (2023b). « FAENet: Frame Averaging Equivariant GNN for Materials Model-
ing. » In: ICML (cit. on pp. 41, 67, 81).
Dwivedi, Vijay Prakash, Ladislav Rampášek, Michael Galkin, Ali Parviz, Guy Wolf, Anh Tuan Luu, and Do-
minique Beaini (2022). « Long range graph benchmark. » In: Advances in Neural Information Processing
Systems 35, pp. 22326–22340 (cit. on p. 19).
Elbayad, Maha, Jiatao Gu, Edouard Grave, and Michael Auli (2020). « Depth-Adaptive Transformer. »
In: International Conference on Learning Representations. URL: [Link]
SJg7KhVKPH (cit. on p. 42).
Eliasof, Moshe, Beatrice Bevilacqua, Carola-Bibiane Schönlieb, and Haggai Maron (2024). « GRANOLA:
Adaptive Normalization for Graph Neural Networks. » In: arXiv preprint arXiv:2404.13344 (cit. on p. 43).
Ennadir, Sofiane, Yassine Abbahaddou, Johannes F. Lutzeyer, Michalis Vazirgiannis, and Henrik Boström
(2024). « A Simple and Yet Fairly Effective Defense for Graph Neural Networks. » In: Proceedings of the
AAAI Conference on Artificial Intelligence. Vol. 38. 19, pp. 21063–21071 (cit. on pp. 5, 56, 76).
Entezari, Negin, Saba A. Al-Sayouri, Amirali Darvishzadeh, and Evangelos E. Papalexakis (2020). « All
You Need Is Low (Rank): Defending Against Adversarial Attacks on Graphs. » In: Proceedings of the
13th International Conference on Web Search and Data Mining, 169–177 (cit. on pp. 56, 64, 67, 68, 76).
Errica, Federico, Henrik Christiansen, Viktor Zaverkin, Takashi Maruyama, Mathias Niepert, and Francesco
Alesiani (2023). « Adaptive Message Passing: A General Framework to Mitigate Oversmoothing, Over-
squashing, and Underreaching. » In: arXiv preprint arXiv:2312.16560 (cit. on p. 43).
Errica, Federico, Marco Podda, Davide Bacciu, and Alessio Micheli (2019). « A fair comparison of graph
neural networks for graph classification. » In: arXiv preprint arXiv:1912.09893 (cit. on p. 90).
– (2020). « A Fair Comparison of Graph Neural Networks for Graph Classification. » In: 8th International
Conference on Learning Representations (cit. on p. 162).
Esser, Pascal, Leena Chennuru Vankadara, and Debarghya Ghoshdastidar (2021). « Learning theory
can (sometimes) explain generalisation in graph neural networks. » In: Advances in Neural Information
Processing Systems 34, pp. 27043–27056 (cit. on p. 81).
Euler, Leonhard (1736). « Solutio problematis ad geometriam situs pertinentis. » In: Commentarii academiae
scientiarum Petropolitanae, pp. 128–140 (cit. on p. 29).
104 BIBLIOGRAPHY

Faber, Lukas and Roger Wattenhofer (2024). « GwAC: GNNs with Asynchronous Communication. » In:
Learning on Graphs Conference. PMLR, pp. 8–1 (cit. on p. 43).
Feng, Fuli, Xiangnan He, Jie Tang, and Tat-Seng Chua (2019). « Graph adversarial training: Dynamically
regularizing based on graph structure. » In: IEEE Transactions on Knowledge and Data Engineering
33.6, pp. 2493–2504 (cit. on p. 55).
Fey, Matthias and Jan E. Lenssen (2019). « Fast Graph Representation Learning with PyTorch Geometric. »
In: ICLR Workshop on Representation Learning on Graphs and Manifolds (cit. on pp. 26, 163, 170, 186).
Fiedler, Miroslav (1973). « Algebraic connectivity of graphs. » In: Czechoslovak mathematical journal 23.2,
pp. 298–305 (cit. on p. 29).
Finkelshtein, Ben, Xingyue Huang, Michael M. Bronstein, and Ismail Ilkan Ceylan (2024). Cooperative
Graph Neural Networks. URL: [Link] (cit. on p. 43).
Freeman, Linton C (1977). « A set of measures of centrality based on betweenness. » In: Sociometry,
pp. 35–41 (cit. on p. 30).
Galkin, Mikhail, Xinyu Yuan, Hesham Mostafa, Jian Tang, and Zhaocheng Zhu (n.d.). « Towards Founda-
tion Models for Knowledge Graph Reasoning. » In: The Twelfth International Conference on Learning
Representations (cit. on p. 3).
Gao, Jian, Jianshe Wu, Xin Zhang, Ying Li, Chunlei Han, and Chubing Guo (2022). « Partition and Learned
Clustering with joined-training: Active learning of GNNs on large-scale graph. » In: Knowledge-Based
Systems 258, p. 110050 (cit. on p. 20).
Gao, Xinbo, Bing Xiao, Dacheng Tao, and Xuelong Li (2010). « A survey of graph edit distance. » In: Pattern
Analysis and applications 13, pp. 113–129 (cit. on p. 21).
Garg, Vikas, Stefanie Jegelka, and Tommi Jaakkola (2020). « Generalization and representational limits
of graph neural networks. » In: International Conference on Machine Learning. PMLR, pp. 3419–3430
(cit. on p. 81).
Gasteiger, Johannes, Aleksandar Bojchevski, and Stephan Günnemann (2019). « Predict then propagate:
Graph neural networks meet personalized pagerank. » In: ICLR (cit. on p. 31).
Gaudelet, Thomas, Ben Day, Arian R Jamasb, Jyothish Soman, Cristian Regep, Gertrude Liu, Jeremy BR
Hayter, Richard Vickers, Charles Roberts, Jian Tang, et al. (2021). « Utilizing graph machine learning
within drug discovery and development. » In: Briefings in bioinformatics 22.6, bbab159 (cit. on p. 81).
Gilmer, Justin, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl (2017). « Neural
message passing for quantum chemistry. » In: International Conference on Machine Learning. PMLR,
pp. 1263–1272 (cit. on pp. 55, 67).
Giraldo, Jhony H., Konstantinos Skianis, Thierry Bouwmans, and Fragkiskos D. Malliaros (2023). « On
the Trade-off between Over-smoothing and Over-squashing in Deep Graph Neural Networks. » In:
Proceedings of the 32nd ACM International Conference on Information and Knowledge Management.
CIKM ’23, 566–576 (cit. on p. 41).
Godwin, Jonathan, Michael Schaarschmidt, Alexander L Gaunt, Alvaro Sanchez-Gonzalez, Yulia Rubanova,
Petar Veličković, James Kirkpatrick, and Peter Battaglia (2022). « Simple GNN Regularisation for 3D
Molecular Property Prediction and Beyond. » In: International Conference on Learning Representations.
URL : [Link] (cit. on p. 21).
Goodfellow, Ian J., Jonathon Shlens, and Christian Szegedy (2015a). « Explaining and Harnessing Ad-
versarial Examples. » In: International Conference on Learning Representations (ICLR) (cit. on pp. 55,
56).
– (2015b). « Explaining and Harnessing Adversarial Examples. » In: International Conference of Learning
Representations (cit. on p. 71).
Goodfellow, Ian, Yoshua Bengio, and Aaron Courville (2016). Deep learning. MIT press (cit. on p. 86).
BIBLIOGRAPHY 105

Gosch, Lukas, Daniel Sturm, Simon Geisler, and Stephan Günnemann (2023). « Revisiting robustness in
graph machine learning. » In: International Conference of Learning Representations (cit. on p. 57).
Graves, Alex (2016). « Adaptive computation time for recurrent neural networks. » In: arXiv preprint
arXiv:1603.08983 (cit. on p. 42).
Grover, Aditya and Jure Leskovec (2016). « node2vec: Scalable feature learning for networks. » In: Pro-
ceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining,
pp. 855–864 (cit. on pp. 16, 37).
Günnemann, Stephan (2022). « Graph neural networks: Adversarial robustness. » In: Graph Neural
Networks: Foundations, Frontiers, and Applications. Springer, pp. 149–176 (cit. on pp. 55, 67, 68).
Guo, Kai, Hongzhi Wen, Wei Jin, Yaming Guo, Jiliang Tang, and Yi Chang (2024). « Investigating out-of-
distribution generalization of GNNs: An architecture perspective. » In: Proceedings of the 30th ACM
SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 932–943 (cit. on pp. 5, 81).
Guo, Kai, Kaixiong Zhou, Xia Hu, Yu Li, Yi Chang, and Xin Wang (2022). « Orthogonal graph neural
networks. » In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 36. 4, pp. 3996–4004
(cit. on p. 60).
Hagberg, Aric, Pieter J Swart, and Daniel A Schult (2008). Exploring network structure, dynamics, and
function using NetworkX. Tech. rep. Los Alamos National Laboratory (LANL), Los Alamos, NM (United
States) (cit. on p. 25).
Hamilton, Will, Zhitao Ying, and Jure Leskovec (2017). « Inductive representation learning on large graphs. »
In: Advances in neural information processing systems 30 (cit. on p. 20).
Hamilton, William L (2020). Graph representation learning. Morgan & Claypool Publishers (cit. on p. 30).
Han, Xiaotian, Zhimeng Jiang, Ninghao Liu, and Xia Hu (2022). « G-mixup: Graph data augmentation for
graph classification. » In: International Conference on Machine Learning. PMLR, pp. 8230–8248 (cit. on
pp. 82, 83, 186).
Härdle, Wolfgang, Axel Werwatz, Marlene Müller, Stefan Sperlich, Wolfgang Härdle, Axel Werwatz, Mar-
lene Müller, and Stefan Sperlich (2004). « Nonparametric density estimation. » In: Nonparametric and
semiparametric models, pp. 39–83 (cit. on p. 182).
Heckerman, David (2008). « A tutorial on learning with Bayesian networks. » In: Innovations in Bayesian
networks: Theory and applications, pp. 33–82 (cit. on p. 69).
Holland, Paul W, Kathryn Blackmond Laskey, and Samuel Leinhardt (1983). « Stochastic blockmodels:
First steps. » In: Social networks 5.2, pp. 109–137 (cit. on p. 35).
Hollocou, Alexandre, Thomas Bonald, and Marc Lelarge (2016). « Improving PageRank for local community
detection. » In: arXiv preprint arXiv:1610.08722 (cit. on p. 16).
Hong, Minui, Jinwoo Choi, and Gunhee Kim (2021). « Stylemix: Separating content and style for enhanced
data augmentation. » In: Proceedings of the IEEE/CVF conference on computer vision and pattern
recognition, pp. 14862–14870 (cit. on p. 82).
Hoseini, Elham, Sattar Hashemi, and Ali Hamzeh (2012). « A levelwise spectral co-clustering algorithm for
collaborative filtering. » In: Proceedings of the 6th international conference on ubiquitous information
management and communication, pp. 1–6 (cit. on p. 13).
Hsu, Chloe, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives
(2022). « Learning inverse folding from millions of predicted structures. » In: International conference on
machine learning. PMLR, pp. 8946–8970 (cit. on p. 21).
Hu, Weihua, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure
Leskovec (2020a). « Open graph benchmark: Datasets for machine learning on graphs. » In: Advances
in neural information processing systems 33, pp. 22118–22133 (cit. on pp. 20, 23, 50, 120).
– (2020b). « Open graph benchmark: Datasets for machine learning on graphs. » In: Advances in Neural
Information Processing Systems 33, pp. 22118–22133 (cit. on p. 62).
106 BIBLIOGRAPHY

Hu, Weihua, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure
Leskovec (2020c). « Open graph benchmark: Datasets for machine learning on graphs. » In: Advances
in neural information processing systems 33, pp. 22118–22133 (cit. on p. 171).
Huang, Gao, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger (2016). « Deep networks with
stochastic depth. » In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The
Netherlands, October 11–14, 2016, Proceedings, Part IV 14. Springer, pp. 646–661 (cit. on p. 41).
Huang, Zheng, Qihui Yang, Dawei Zhou, and Yujun Yan (2024). « Enhancing Size Generalization in Graph
Neural Networks through Disentangled Representation Learning. » In: arXiv preprint arXiv:2406.04601
(cit. on p. 81).
Huré, Côme, Huyên Pham, Achref Bachouch, and Nicolas Langrené (2021). « Deep neural networks
algorithms for stochastic control problems on finite horizon: convergence analysis. » In: SIAM Journal on
Numerical Analysis 59.1, pp. 525–557 (cit. on p. 48).
Ivanov, Sergei, Sergei Sviridov, and Evgeny Burnaev (2019). « Understanding isomorphism bias in graph
data sets. » In: arXiv preprint arXiv:1910.12091 (cit. on p. 24).
Jacot, Arthur, Franck Gabriel, and Clément Hongler (2018). « Neural tangent kernel: Convergence and
generalization in neural networks. » In: Advances in neural information processing systems 31 (cit. on
pp. 19, 81).
Jagtap, Surabhi, Abdulkadir Çelikkanat, Aurélie Pirayre, Frédérique Bidard, Laurent Duval, and Fragkiskos
D. Malliaros (2022). « BraneMF: integration of biological networks for functional analysis of proteins. » In:
Bioinform. 38.24, pp. 5383–5389 (cit. on p. 81).
Jia, Xu, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool (2016). « Dynamic filter networks. » In:
Advances in neural information processing systems 29 (cit. on p. 43).
Jin, Ming, Heng Chang, Wenwu Zhu, and Somayeh Sojoudi (2021). « Power up! robust graph convolutional
network via graph powering. » In: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 8004–
8012 (cit. on p. 31).
Joshi, Chaitanya K, Cristian Bodnar, Simon V Mathis, Taco Cohen, and Pietro Lio (2023). « On the
expressive power of geometric graph neural networks. » In: International conference on machine learning.
PMLR, pp. 15330–15355 (cit. on p. 20).
Jovanović, Irena and Zoran Stanić (2012). « Spectral distances of graphs. » In: Linear algebra and its
applications 436.5, pp. 1425–1435 (cit. on p. 21).
Ju, Mingxuan, Yujie Fan, Chuxu Zhang, and Yanfang Ye (2023). « Let graph be the go board: gradient-free
node injection attack for graph neural networks via reinforcement learning. » In: Proceedings of the AAAI
Conference on Artificial Intelligence. Vol. 37. 4, pp. 4383–4390 (cit. on p. 56).
Kaur, Arshleen, Vinay Kukreja, and Amanveer Singh (2024). « Molecular Insights Unveiled: A Hybrid Neural
Network Model GNN-CNN for Drug Discovery. » In: 2024 International Conference on Big Data Analytics
in Bioinformatics (DABCon). IEEE, pp. 1–5 (cit. on p. 97).
Kearnes, Steven, Kevin McCloskey, Marc Berndl, Vijay Pande, and Patrick Riley (2016). « Molecular
graph convolutions: moving beyond fingerprints. » In: Journal of Computer-Aided Molecular Design 30.8,
pp. 595–608 (cit. on p. 55).
Kingma, Diederik P and Jimmy Lei Ba (2015a). « Adam: A method for stochastic optimization. » In:
International Conference of Learning Representations (cit. on p. 76).
Kingma, Diederik P. and Jimmy Ba (2015b). « Adam: A Method for Stochastic Optimization. » In: Interna-
tional Conference on Learning Representations (cit. on pp. 50, 120, 121, 162, 186).
Kipf, Thomas N. and Max Welling (2017a). « Semi-Supervised Classification with Graph Convolutional
Networks. » In: ICLR (cit. on pp. 3, 17, 41, 55, 67, 76).
– (2017b). « Semi-Supervised Classification with Graph Convolutional Networks. » In: International Confer-
ence on Learning Representations (cit. on pp. 31, 33, 38).
BIBLIOGRAPHY 107

– (2017c). « Semi-Supervised Classification with Graph Convolutional Networks. » In: International Confer-
ence on Learning Representations (cit. on p. 179).
Koh, Pang Wei and Percy Liang (2017). « Understanding black-box predictions via influence functions. » In:
International conference on machine learning. PMLR, pp. 1885–1894 (cit. on pp. 88, 187).
Koke, Christian and Daniel Cremers (2024). « HoloNets: Spectral Convolutions do extend to Directed
Graphs. » In: International Conference on Learning Representations (ICLR) (cit. on p. 31).
Kong, Shuming, Yanyan Shen, and Linpeng Huang (2021). « Resolving training biases via influence-based
data relabeling. » In: International Conference on Learning Representations (cit. on p. 88).
König, Dénes (1936). « Theorie der endlichen und unendlichen Graphen: Kombinatorische Topologie der
Streckenkomplexe. » In: Akademische Verlagsgesellschaft M. R. H., Leipzig (cit. on p. 29).
Kowalski, Emmanuel (2019). An introduction to expander graphs. Société mathématique de France Paris
(cit. on p. 11).
Kreuzer, Devin, Dominique Beaini, Will Hamilton, Vincent Létourneau, and Prudencio Tossou (2021).
« Rethinking graph transformers with spectral attention. » In: Advances in Neural Information Processing
Systems 34, pp. 21618–21629 (cit. on p. 31).
Krishnan, Vishaal, Al Makdah, Abed AlRahman, and Fabio Pasqualetti (2020). « Lipschitz bounds and
provably robust training by laplacian smoothing. » In: Advances in Neural Information Processing Systems
33, pp. 10924–10935 (cit. on p. 69).
Krizhevsky, A., I. Sutskever, and G. E. Hinton (2012). « ImageNet classification with deep convolutional
neural networks. » In: Advances in Neural Information Processing Systems (NeurIPS) (cit. on p. 82).
Lafferty, John D., Andrew McCallum, and Fernando C. N. Pereira (2001). « Conditional Random Fields:
Probabilistic Models for Segmenting and Labeling Sequence Data. » In: Proceedings of the Eighteenth
International Conference on Machine Learning (ICML), pp. 282–289 (cit. on p. 68).
Lange, Oliver and Luis Perez (2020). « Traffic prediction with advanced graph neural networks. » In:
DeepMind Research Blog Post (cit. on p. 97).
Law, John (1986). Robust statistics—the approach based on influence functions (cit. on p. 88).
Lee, Byung-Kwan, Junho Kim, and Yong Man Ro (2022). « Masking Adversarial Damage: Finding Ad-
versarial Saliency for Robust and Sparse Network. » In: Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition (CVPR), pp. 15126–15136 (cit. on p. 88).
Lee, Clement and Darren J Wilkinson (2019). « A review of stochastic block models and extensions for
graph clustering. » In: Applied Network Science 4.1, pp. 1–50 (cit. on p. 23).
Lee, John Boaz, Ryan A Rossi, Xiangnan Kong, Sungchul Kim, Eunyee Koh, and Anup Rao (2019). « Graph
convolutional networks with motif-based attention. » In: Proceedings of the 28th ACM international
conference on information and knowledge management, pp. 499–508 (cit. on p. 31).
Lei, Runlin, Zhen Wang, Yaliang Li, Bolin Ding, and Zhewei Wei (2022). « Evennet: Ignoring odd-hop neigh-
bors improves robustness of graph neural networks. » In: Advances in Neural Information Processing
Systems 35, pp. 4694–4706 (cit. on p. 160).
Leskovec, Jure, Deepayan Chakrabarti, Jon Kleinberg, Christos Faloutsos, and Zoubin Ghahramani (2010).
« Kronecker graphs: an approach to modeling networks. » In: Journal of Machine Learning Research
11.2 (cit. on p. 33).
Leskovec, Jure and Julian Mcauley (2012). « Learning to discover social circles in ego networks. » In:
Advances in neural information processing systems 25 (cit. on p. 33).
Li, Hao, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf (2017). « Pruning Filters
for Efficient ConvNets. » In: International Conference on Learning Representations. URL: https : / /
[Link]/forum?id=rJqFGTslg (cit. on p. 41).
108 BIBLIOGRAPHY

Li, Haoyang, Xin Wang, Ziwei Zhang, and Wenwu Zhu (2022). « Ood-gnn: Out-of-distribution generalized
graph neural network. » In: IEEE Transactions on Knowledge and Data Engineering 35.7, pp. 7328–7340
(cit. on pp. 5, 19, 81).
Li, Kuan, YiWen Chen, Yang Liu, Jin Wang, Qing He, Minhao Cheng, and Xiang Ao (2024a). « Boosting the
Adversarial Robustness of Graph Neural Networks: An OOD Perspective. » In: The Twelfth International
Conference on Learning Representations (cit. on p. 76).
Li, Pan and Jure Leskovec (2022). « The expressive power of graph neural networks. » In: Graph Neural
Networks: Foundations, Frontiers, and Applications, pp. 63–98 (cit. on p. 20).
Li, Rui, Chaozhuo Li, Yanming Shen, Zeyu Zhang, and Xu Chen (2024b). « Generalizing Knowledge Graph
Embedding with Universal Orthogonal Parameterization. » In: arXiv preprint arXiv:2405.08540 (cit. on
p. 8).
Liao, Renjie, Raquel Urtasun, and Richard Zemel (2020). « A pac-bayesian approach to generalization
bounds for graph neural networks. » In: arXiv preprint arXiv:2012.07690 (cit. on pp. 19, 81).
Lim, Derek and Austin R Benson (2021). « Expertise and dynamics within crowdsourced musical knowledge
curation: A case study of the genius platform. » In: Proceedings of the International AAAI Conference on
Web and Social Media. Vol. 15, pp. 373–384 (cit. on pp. 23, 50, 120).
Lim, Derek, Felix Matthew Hohne, Xiuyu Li, Sijia Linda Huang, Vaishnavi Gupta, Omkar Prasad Bhalerao,
and Ser-Nam Lim (2021a). « Large Scale Learning on Non-Homophilous Graphs: New Benchmarks and
Strong Simple Methods. » In: Advances in Neural Information Processing Systems. Ed. by A. Beygelzimer,
Y. Dauphin, P. Liang, and J. Wortman Vaughan. URL: [Link]
(cit. on pp. 23, 50, 120, 121).
Lim, Derek, Xiuyu Li, Felix Hohne, and Ser-Nam Lim (2021b). « New Benchmarks for Learning on Non-
Homophilous Graphs. » In: arXiv preprint arXiv:2104.01404 (cit. on p. 170).
Ling, Hongyi, Zhimeng Jiang, Meng Liu, Shuiwang Ji, and Na Zou (2023). « Graph mixup with soft
alignments. » In: International Conference on Machine Learning. PMLR, pp. 21335–21349 (cit. on pp. 82,
83).
Liu, Gary, Denise B Catacutan, Khushi Rathod, Kyle Swanson, Wengong Jin, Jody C Mohammed, Anush
Chiappino-Pepe, Saad A Syed, Meghan Fragis, Kenneth Rachwalski, et al. (2023). « Deep learning-
guided discovery of an antibiotic targeting Acinetobacter baumannii. » In: Nature Chemical Biology 19.11,
pp. 1342–1350 (cit. on p. 96).
Liu, Juncheng, Kenji Kawaguchi, Bryan Hooi, Yiwei Wang, and Xiaokui Xiao (2021a). « Eignn: Efficient
infinite-depth graph neural networks. » In: Advances in Neural Information Processing Systems 34,
pp. 18762–18773 (cit. on p. 41).
Liu, Xiaorui, Jiayuan Ding, Wei Jin, Han Xu, Yao Ma, Zitao Liu, and Jiliang Tang (2021b). « Graph neural
networks with adaptive residual. » In: Advances in Neural Information Processing Systems 34, pp. 9720–
9733 (cit. on pp. 57, 63, 68).
Liu, Xiaorui, Wei Jin, Yao Ma, Yaxin Li, Hua Liu, Yiqi Wang, Ming Yan, and Jiliang Tang (2021c). « Elastic
Graph Neural Networks. » In: International Conference on Machine Learning, pp. 6837–6849 (cit. on
p. 160).
Liu, Yangwei, Hu Ding, Danyang Chen, and Jinhui Xu (2017). « Novel geometric approach for global
alignment of PPI networks. » In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 31. 1
(cit. on p. 3).
Liu, Yazheng and Sihong Xie (n.d.). « Explanations of GNN on Evolving Graphs via Axiomatic Layer
edges. » In: The Thirteenth International Conference on Learning Representations (cit. on p. 97).
Lovász, László (1993). « Random walks on graphs. » In: Combinatorics, Paul erdos is eighty 2.1-46, p. 4
(cit. on p. 11).
BIBLIOGRAPHY 109

Luan, Sitao, Chenqing Hua, Qincheng Lu, Jiaqi Zhu, Mingde Zhao, Shuyuan Zhang, Xiao-Wen Chang,
and Doina Precup (2022). « Revisiting heterophily for graph neural networks. » In: Advances in neural
information processing systems 35, pp. 1362–1375 (cit. on p. 41).
Luan, Sitao, Chenqing Hua, Minkai Xu, Qincheng Lu, Jiaqi Zhu, Xiao-Wen Chang, Jie Fu, Jure Leskovec,
and Doina Precup (2024). « When do graph neural networks help with node classification? investigating
the homophily principle on node distinguishability. » In: Advances in Neural Information Processing
Systems 36 (cit. on p. 8).
Lutzeyer, Johannes (2020). Network Representation Matrices and their Eigenproperties: A Comparative
Study. PhD Thesis: Imperial College London (cit. on p. 35).
Ma, Liheng, Chen Lin, Derek Lim, Adriana Romero-Soriano, Puneet K Dokania, Mark Coates, Philip Torr,
and Ser-Nam Lim (2023). « Graph inductive biases in transformers without message passing. » In:
International Conference on Machine Learning. PMLR, pp. 23321–23337 (cit. on p. 31).
Ma, Xiaoxiao, Jia Wu, Shan Xue, Jian Yang, Chuan Zhou, Quan Z Sheng, Hui Xiong, and Leman Akoglu
(2021). « A comprehensive survey on graph anomaly detection with deep learning. » In: IEEE Transactions
on Knowledge and Data Engineering (cit. on p. 56).
Ma, Xinyu, Xu Chu, Yasha Wang, Yang Lin, Junfeng Zhao, Liantao Ma, and Wenwu Zhu (2024). « Fused
gromov-wasserstein graph mixup for graph-level classifications. » In: Advances in Neural Information
Processing Systems 36 (cit. on p. 83).
Malliaros, Fragkiskos D., Christos Giatsidis, Apostolos N. Papadopoulos, and Michalis Vazirgiannis (2020).
« The core decomposition of networks: theory, algorithms and applications. » In: VLDB J. 29.1, pp. 61–92
(cit. on pp. 14, 29, 32, 48).
Malliaros, Fragkiskos D, Maria-Evgenia G Rossi, and Michalis Vazirgiannis (2016). « Locating influential
nodes in complex networks. » In: Scientific reports 6.1, p. 19307 (cit. on p. 14).
Malliaros, Fragkiskos D and Michalis Vazirgiannis (2013a). « Clustering and community detection in directed
networks: A survey. » In: Physics reports 533.4, pp. 95–142 (cit. on pp. 3, 10).
Malliaros, Fragkiskos D. and Michalis Vazirgiannis (2013b). « Clustering and community detection in
directed networks: A survey. » In: Physics Reports 533.4. Clustering and Community Detection in
Directed Networks: A Survey, pp. 95–142 (cit. on p. 81).
Mateos, Gonzalo, Santiago Segarra, Antonio G Marques, and Alejandro Ribeiro (2019). « Connecting the
dots: Identifying network structure via graph signal processing. » In: IEEE Signal Processing Magazine
36.3, pp. 16–43 (cit. on p. 12).
Minaee, Shervin, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain,
and Jianfeng Gao (2024). « Large language models: A survey. » In: arXiv preprint arXiv:2402.06196
(cit. on p. 67).
Modell, Alexander and Patrick Rubin-Delanchy (2021). « Spectral clustering under degree heterogeneity: a
case for the random walk Laplacian. » In: arXiv preprint arXiv:2105.00987 (cit. on pp. 11, 12).
Mohapatra, Jeet, Ching-Yun Ko, Lily Weng, Pin-Yu Chen, Sijia Liu, and Luca Daniel (2021). « Hidden cost
of randomized smoothing. » In: International Conference on Artificial Intelligence and Statistics. PMLR,
pp. 4033–4041 (cit. on p. 69).
Morris, Christopher, Nils M Kriege, Franka Bause, Kristian Kersting, Petra Mutzel, and Marion Neumann
(2020a). « TUDataset: A collection of benchmark datasets for learning with graphs. » In: Graph Repre-
sentation Learning and Beyond (GRL+), ICML Workshop (cit. on pp. 62, 162).
Morris, Christopher, Nils M. Kriege, Franka Bause, Kristian Kersting, Petra Mutzel, and Marion Neumann
(2020b). « TUDataset: A collection of benchmark datasets for learning with graphs. » In: ICML 2020
Workshop on Graph Representation Learning and Beyond (GRL+ 2020). arXiv: 2007.08663 (cit. on
p. 186).
110 BIBLIOGRAPHY

Morris, Christopher, Martin Ritzert, Matthias Fey, William L Hamilton, Jan Eric Lenssen, Gaurav Rattan,
and Martin Grohe (2019). « Weisfeiler and leman go neural: Higher-order graph neural networks. » In:
Proceedings of the AAAI conference on artificial intelligence. Vol. 33. 01, pp. 4602–4609 (cit. on p. 20).
Nagisa, Masaru (2022). « The p-norm of some matrices. » In: arXiv preprint arXiv:2209.08553 (cit. on
p. 158).
Naim, Iftekhar and Daniel Gildea (2012). « Convergence of the EM algorithm for Gaussian mixtures with
unbalanced mixing coefficients. » In: arXiv preprint arXiv:1206.6427 (cit. on p. 186).
Nelsen, Roger B (2006). An introduction to copulas. Springer (cit. on p. 182).
Newman, Mark EJ, Duncan J Watts, and Steven H Strogatz (2002). « Random graph models of social
networks. » In: Proceedings of the national academy of sciences 99.suppl_1, pp. 2566–2572 (cit. on
p. 81).
Newman, Mark (2013). Networks: An Introduction (cit. on pp. 91, 181).
Ng, Andrew (2000). « CS229 Lecture notes. » In: CS229 Lecture notes 1.1, pp. 1–3 (cit. on p. 87).
Ng, Andrew, Michael Jordan, and Yair Weiss (2001). « On spectral clustering: Analysis and an algorithm. »
In: Advances in neural information processing systems 14 (cit. on p. 35).
Nielsen, Frank (2013). « Cramér-Rao lower bound and information geometry. » In: Connected at Infinity II:
A Selection of Mathematics by Indians, pp. 18–37 (cit. on p. 88).
Nikolentzos, Giannis, George Dasoulas, and Michalis Vazirgiannis (2020). « k-hop graph neural networks. »
In: Neural Networks 130, pp. 195–205 (cit. on p. 31).
Nikolentzos, Giannis and Michalis Vazirgiannis (2020). « Random walk graph neural networks. » In:
Advances in Neural Information Processing Systems 33, pp. 16211–16222 (cit. on p. 21).
Ortega, Antonio, Pascal Frossard, Jelena Kovačević, José MF Moura, and Pierre Vandergheynst (2018).
« Graph signal processing: Overview, challenges, and applications. » In: Proceedings of the IEEE 106.5,
pp. 808–828 (cit. on p. 29).
Ozertem, Umut and Deniz Erdogmus (2011). « Locally defined principal curves and surfaces. » In: The
Journal of Machine Learning Research 12, pp. 1249–1286 (cit. on p. 186).
Page, Lawrence (1999). The PageRank citation ranking: Bringing order to the web. Tech. rep. Technical
Report (cit. on p. 11).
Paszke, Adam, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor
Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. (2019). « Pytorch: An imperative style,
high-performance deep learning library. » In: Advances in neural information processing systems 32
(cit. on p. 25).
Peng, Chengbin, Tamara G Kolda, and Ali Pinar (2014). « Accelerating community detection by using
k-core subgraphs. » In: arXiv preprint arXiv:1403.2226 (cit. on p. 15).
Perozzi, Bryan, Rami Al-Rfou, and Steven Skiena (2014). « Deepwalk: Online learning of social representa-
tions. » In: Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and
data mining, pp. 701–710 (cit. on p. 16).
Pfaff, Tobias, Meire Fortunato, Alvaro Sanchez-Gonzalez, and Peter W Battaglia (2020). « Learning
mesh-based simulation with graph networks. » In: arXiv preprint arXiv:2010.03409 (cit. on pp. 19, 81).
Platonov, Oleg, Denis Kuznedelev, Artem Babenko, and Liudmila Prokhorenkova (2024). « Characterizing
graph datasets for node classification: Homophily-heterophily dichotomy and beyond. » In: Advances in
Neural Information Processing Systems 36 (cit. on p. 8).
Poincaré, Henri (1900). « Second complément à l’analysis situs. » In: Proceedings of the London Mathe-
matical Society 1.1, pp. 277–308 (cit. on p. 29).
Pons, Pascal and Matthieu Latapy (2005). « Computing communities in large networks using random
walks. » In: Computer and Information Sciences-ISCIS 2005: 20th International Symposium, Istanbul,
Turkey, October 26-28, 2005. Proceedings 20. Springer, pp. 284–293 (cit. on p. 37).
BIBLIOGRAPHY 111

Raıssouli, Mustapha and Iqbal H Jebril (2010). « Various Proofs for the Decrease Monotonicity of the
Schatten’s Power Norm, Various Families of R n- Norms and Some Open Problems. » In: Int. J. Open
Problems Compt. Math 3.2, pp. 164–174 (cit. on p. 148).
Rampášek, Ladislav, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique
Beaini (2022a). « Recipe for a general, powerful, scalable graph transformer. » In: Advances in Neural
Information Processing Systems 35, pp. 14501–14515 (cit. on p. 31).
– (2022b). « Recipe for a general, powerful, scalable graph transformer. » In: Advances in Neural Informa-
tion Processing Systems 35, pp. 14501–14515 (cit. on p. 41).
Rebuffi, Sylvestre-Alvise, Sven Gowal, Dan Andrei Calian, Florian Stimberg, Olivia Wiles, and Timothy
A Mann (2021). « Data augmentation can improve robustness. » In: Advances in Neural Information
Processing Systems 34, pp. 29935–29948 (cit. on p. 82).
Rice, Leslie, Anna Bair, Huan Zhang, and J Zico Kolter (2021a). « Robustness between the worst and
average case. » In: Advances in Neural Information Processing Systems 34, pp. 27840–27851 (cit. on
p. 58).
– (2021b). « Robustness between the worst and average case. » In: Advances in Neural Information
Processing Systems 34, pp. 27840–27851 (cit. on p. 75).
Rong, Yu, Wenbing Huang, Tingyang Xu, and Junzhou Huang (2019). « Dropedge: Towards deep graph
convolutional networks on node classification. » In: arXiv preprint arXiv:1907.10903 (cit. on pp. 82, 186).
Rozemberczki, Benedek, Carl Allen, and Rik Sarkar (2021). « Multi-scale attributed node embedding. » In:
Journal of Complex Networks 9.2, cnab014 (cit. on pp. 23, 50, 120, 121).
Rozemberczki, Benedek and Rik Sarkar (2020). « Characteristic functions on graphs: Birds of a feather,
from statistical descriptors to parametric models. » In: Proceedings of the 29th ACM international
conference on information & knowledge management, pp. 1325–1334 (cit. on pp. 23, 120).
Sabour, Sara, Nicholas Frosst, and Geoffrey E Hinton (2017). « Dynamic routing between capsules. » In:
Advances in neural information processing systems 30 (cit. on p. 41).
Salha-Galvan, Guillaume (2022). « Contributions to Representation Learning with Graph Autoencoders
and Applications to Music Recommendation. » In: arXiv preprint arXiv:2205.14651 (cit. on p. 96).
Salimans, Tim and Durk P Kingma (2016). « Weight normalization: A simple reparameterization to accel-
erate training of deep neural networks. » In: Advances in Neural Information Processing Systems 29
(cit. on p. 61).
Sandryhaila, Aliaksei and José MF Moura (2013). « Discrete signal processing on graphs. » In: IEEE
transactions on signal processing 61.7, pp. 1644–1656 (cit. on p. 29).
Sarkar, Camellia and Sarika Jalan (2018). « Spectral properties of complex networks. » In: Chaos: An
Interdisciplinary Journal of Nonlinear Science 28.10 (cit. on p. 11).
Schichl, Hermann and Meinolf Sellmann (2015). « Predisaster preparation of transportation networks. » In:
Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 29. 1 (cit. on p. 3).
Scholten, Yan, Jan Schuchardt, Simon Geisler, Aleksandar Bojchevski, and Stephan Günnemann (2022).
« Randomized message-interception smoothing: Gray-box certificates for graph neural networks. » In:
Advances in Neural Information Processing Systems 35, pp. 33146–33158 (cit. on p. 57).
Schuchardt, Jan, Aleksandar Bojchevski, Johannes Gasteiger, and Stephan Günnemann (2021). « Collec-
tive Robustness Certificates: Exploiting Interdependence in Graph Neural Networks. » In: International
Conference on Learning Representations (cit. on p. 55).
Schuster, Tal, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Don-
ald Metzler (2022). « Confident Adaptive Language Modeling. » In: Advances in Neural Information
Processing Systems. Ed. by Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho. URL:
[Link] (cit. on p. 42).
112 BIBLIOGRAPHY

Seddik, Mohamed El Amine, Changmin Wu, Johannes F Lutzeyer, and Michalis Vazirgiannis (2022).
« Node feature kernels increase graph convolutional network robustness. » In: International Conference
on Artificial Intelligence and Statistics. PMLR, pp. 6225–6241 (cit. on pp. 57, 63, 69).
Seidman, Stephen B (1983). « Network structure and minimum degree. » In: Social networks 5.3, pp. 269–
287 (cit. on p. 29).
Sen, Prithviraj, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad (2008).
« Collective Classification in Network Data. » In: AI Magazine 29.3, p. 93. DOI: 10.1609/aimag.v29i3.
2157. URL: [Link] (cit. on pp. 23, 36, 50,
62, 120, 161, 170).
Shalev-Shwartz, Shai and Shai Ben-David (2014). Understanding machine learning: From theory to
algorithms. Cambridge university press (cit. on p. 84).
Shchur, Oleksandr, Maximilian Mumme, Aleksandar Bojchevski, and Stephan Günnemann (2018a). « Pit-
falls of graph neural network evaluation. » In: arXiv preprint arXiv:1811.05868 (cit. on p. 20).
– (2018b). « Pitfalls of graph neural network evaluation. » In: Relational Representation Learning Workshop
(R2L 2018), NeurIPS (cit. on pp. 23, 44, 50, 62, 120, 161, 170).
Shervashidze, Nino, Pascal Schweitzer, Erik Jan Van Leeuwen, Kurt Mehlhorn, and Karsten M Borgwardt
(2011). « Weisfeiler-lehman graph kernels. » In: Journal of Machine Learning Research 12.9 (cit. on
p. 21).
Shim, Kyuhong, Jungwook Choi, and Wonyong Sung (2021). « Understanding the role of self attention for
efficient speech recognition. » In: International Conference on Learning Representations (cit. on p. 67).
Son, Jieun and Seoung Bum Kim (2017). « Content-based filtering for recommendation systems using
multiattribute networks. » In: Expert Systems with Applications 89, pp. 404–412 (cit. on pp. 3, 14).
Stokes, Jonathan M, Kevin Yang, Kyle Swanson, Wengong Jin, Andres Cubillos-Ruiz, Nina M Donghia,
Craig R MacNair, Shawn French, Lindsey A Carfrae, Zohar Bloom-Ackermann, et al. (2020). « A deep
learning approach to antibiotic discovery. » In: Cell 180.4, pp. 688–702 (cit. on p. 3).
Sun, Ke, Zhanxing Zhu, and Zhouchen Lin (2019). « Adagcn: Adaboosting graph convolutional networks
into deep models. » In: ICLR (cit. on pp. 31, 50).
Sutton, Charles, Andrew McCallum, et al. (2012). « An introduction to conditional random fields. » In:
Foundations and Trends® in Machine Learning 4.4, pp. 267–373 (cit. on p. 69).
Tang, Huayi and Yong Liu (2023). « Towards understanding generalization of graph neural networks. » In:
International Conference on Machine Learning. PMLR, pp. 33674–33719 (cit. on pp. 19, 81).
Tang, Jian, Meng Qu, Mingzhe Wang, Ming Zhang, Jun Yan, and Qiaozhu Mei (2015). « Line: Large-scale
information network embedding. » In: Proceedings of the 24th international conference on world wide
web, pp. 1067–1077 (cit. on p. 16).
Tang, Xianfeng, Yandong Li, Yiwei Sun, Huaxiu Yao, Prasenjit Mitra, and Suhang Wang (2020a). « Trans-
ferring robustness for graph neural network against poisoning attacks. » In: Proceedings of the 13th
international Conference on Web Search and Data Mining, pp. 600–608 (cit. on pp. 56, 68).
– (2020b). « Transferring Robustness for Graph Neural Network Against Poisoning Attacks. » In: ACM
Internatioal Conference on Web Search and Data Mining (WSDM) (cit. on p. 68).
Tao, Shuchang, Qi Cao, Huawei Shen, Junjie Huang, Yunfan Wu, and Xueqi Cheng (2021). « Single
Node Injection Attack against Graph Neural Networks. » In: Proceedings of the 30th ACM International
Conference on Information & Knowledge Management. ACM (cit. on p. 56).
Traud, Amanda L, Peter J Mucha, and Mason A Porter (2012). « Social structure of facebook networks. »
In: Physica A: Statistical Mechanics and its Applications 391.16, pp. 4165–4180 (cit. on pp. 23, 120).
Tsitsulin, Anton, Benedek Rozemberczki, John Palowitch, and Bryan Perozzi (2022). « Synthetic graph
generation to benchmark graph learning. » In: arXiv preprint arXiv:2204.01376 (cit. on p. 22).
BIBLIOGRAPHY 113

Tzikas, Dimitris G, Aristidis C Likas, and Nikolaos P Galatsanos (2008). « The variational approximation for
Bayesian inference. » In: IEEE Signal Processing Magazine 25.6, pp. 131–146 (cit. on p. 182).
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz
Kaiser, and Illia Polosukhin (2017). « Attention Is All You Need. » In: CoRR abs/1706.03762. arXiv:
1706.03762. URL: [Link] (cit. on p. 43).
Vela, Ariel Ricardo Ramos, Johannes F Lutzeyer, Anastasios Giovanidis, and Michalis Vazirgiannis (n.d.).
« Improving Graph Neural Networks at Scale: Combining Approximate PageRank and CoreRank. » In:
NeurIPS 2022 Workshop: New Frontiers in Graph Learning (cit. on p. 31).
Veličković, Petar, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio
(2018). « Graph Attention Networks. » In: International Conference on Learning Representations. URL:
[Link] (cit. on pp. 17, 38).
Velickovic, Petar, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio
(2018). « Graph Attention Networks. » In: 6th International Conference on Learning Representations,
(ICLR) (cit. on p. 67).
Vignac, Clement, Igor Krawczuk, Antoine Siraudin, Bohan Wang, Volkan Cevher, and Pascal Frossard
(2023). « DiGress: Discrete Denoising diffusion for graph generation. » In: The Eleventh International
Conference on Learning Representations (cit. on p. 81).
Von Luxburg, Ulrike (2007). « A tutorial on spectral clustering. » In: Statistics and computing 17, pp. 395–
416 (cit. on pp. 12, 14, 35).
Wahl, Scott and John Sheppard (2015). « Hierarchical fuzzy spectral clustering in social networks using
spectral characterization. » In: The twenty-eighth international flairs conference (cit. on p. 12).
Wallach, Hanna M et al. (2004). « Conditional random fields: An introduction. » In: University of Pennsylvania
CIS Technical Report MS-CIS-04-21 24, pp. 33–42 (cit. on p. 69).
Wang, Hongwei and Jure Leskovec (2020). « Unifying graph convolutional neural networks and label
propagation. » In: arXiv preprint arXiv:2002.06755 (cit. on p. 8).
Wang, Hongwei, Miao Zhao, Xing Xie, Wenjie Li, and Minyi Guo (2019). « Knowledge graph convolutional
networks for recommender systems. » In: The world wide web conference, pp. 3307–3313 (cit. on p. 3).
Wang, Lei, Runtian Zhai, Di He, Liwei Wang, and Li Jian (2021a). « Pretrain-to-finetune adversarial training
via sample-wise randomized smoothing. » In: Openreview (cit. on p. 69).
Wang, Yiwei, Shenghua Liu, Minji Yoon, Hemank Lamba, Wei Wang, Christos Faloutsos, and Bryan
Hooi (2020a). « Provably robust node classification via low-pass message passing. » In: 2020 IEEE
International Conference on Data Mining (ICDM). IEEE, pp. 621–630 (cit. on p. 56).
Wang, Yiwei, Wei Wang, Yuxuan Liang, Yujun Cai, and Bryan Hooi (2021b). « Mixup for node and graph
classification. » In: Proceedings of the Web Conference 2021, pp. 3663–3674 (cit. on p. 83).
Wang, Yulin, Kangchen Lv, Rui Huang, Shiji Song, Le Yang, and Gao Huang (2020b). « Glance and
Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification. » In: CoRR
abs/2010.05300. arXiv: 2010.05300. URL: [Link] (cit. on p. 42).
Weng, Tsui-Wei, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, Dong Su, Yupeng Gao, Cho-Jui Hsieh, and Luca
Daniel (2018a). « Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach. »
In: International Conference on Learning Representations (ICLR) (cit. on p. 61).
– (2018b). « Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach. » In:
International Conference on Learning Representations (cit. on p. 71).
Wilkinson, Darren J (2018). Stochastic modelling for systems biology. Chapman and Hall/CRC (cit. on
p. 23).
Wu, Felix, Amauri Souza, Tianyi Zhang, Christopher Fifty, Tao Yu, and Kilian Weinberger (2019a). « Simpli-
fying graph convolutional networks. » In: International conference on machine learning. PMLR, pp. 6861–
6871 (cit. on p. 38).
114 BIBLIOGRAPHY

Wu, Huijun, Chen Wang, Yuriy Tyshetskiy, Andrew Docherty, Kai Lu, and Liming Zhu (2019b). « Adversarial
Examples for Graph Data: Deep Insights into Attack and Defense. » In: Proceedings of the Twenty-Eighth
International Joint Conference on Artificial Intelligence, IJCAI-19, pp. 4816–4823 (cit. on pp. 56, 64, 67,
68, 76).
Wu, Jing, Mingyi Zhou, Ce Zhu, Yipeng Liu, Mehrtash Harandi, and Li Li (2021). « Performance evaluation
of adversarial attacks: Discrepancies and solutions. » In: arXiv preprint arXiv:2104.11103 (cit. on p. 71).
Wu, Shu, Yuyuan Tang, Yanqiao Zhu, Liang Wang, Xing Xie, and Tieniu Tan (2019c). « Session-based
Recommendation with Graph Neural Networks. » In: Proceedings of the 33rd AAAI Conference on
Artificial Intelligence, pp. 346–353 (cit. on p. 67).
Wu, Zhenqin, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu,
Karl Leswing, and Vijay Pande (2018). « MoleculeNet: a benchmark for molecular machine learning. » In:
Chemical science 9.2, pp. 513–530 (cit. on p. 20).
Xu, Kaidi, Hongge Chen, Sijia Liu, Pin-Yu Chen, Tsui-Wei Weng, Mingyi Hong, and Xue Lin (2019a).
« Topology attack and defense for graph neural networks: An optimization perspective. » In: Proceedings
of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI-19) (cit. on pp. 61, 62,
64, 68, 76, 161).
Xu, Keyulu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka (2019b). « How Powerful are Graph Neural
Networks? » In: International Conference on Learning Representations. URL: [Link]
forum?id=ryGs6iA5Km (cit. on pp. 3, 18, 20, 32, 38, 180).
– (2019c). « How Powerful are Graph Neural Networks? » In: 7th International Conference on Learning
Representations (cit. on pp. 41, 55, 67).
Xu, Keyulu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka
(2018b). « Representation Learning on Graphs with Jumping Knowledge Networks. » In: Proceedings of
the 35th International Conference on Machine Learning. Ed. by Jennifer Dy and Andreas Krause. Vol. 80.
Proceedings of Machine Learning Research. PMLR, pp. 5453–5462. URL: [Link]
press/v80/[Link] (cit. on p. 50).
Xu, Keyulu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie
Jegelka (2018a). « Representation Learning on Graphs with Jumping Knowledge Networks. » In: CoRR
abs/1806.03536. arXiv: 1806.03536. URL: [Link] (cit. on p. 43).
Yang, Chenxiao, Qitian Wu, Jiahua Wang, and Junchi Yan (2022). « Graph neural networks are inherently
good generalizers: Insights by bridging gnns and mlps. » In: arXiv preprint arXiv:2212.09034 (cit. on
p. 81).
Yang, Liang, Runjie Shi, Qiuliang Zhang, Bingxin Niu, Zhen Wang, Xiaochun Cao, and Chuan Wang (2023).
« Self-supervised Graph Neural Networks via Low-Rank Decomposition. » In: Thirty-seventh Conference
on Neural Information Processing Systems (cit. on p. 69).
Yang, Miin-Shen, Chien-Yo Lai, and Chih-Ying Lin (2012). « A robust EM clustering algorithm for Gaussian
mixture models. » In: Pattern Recognition 45.11, pp. 3950–3961 (cit. on pp. 86, 87, 182).
Yang, Renchi, Jieming Shi, Xiaokui Xiao, Yin Yang, and Sourav S Bhowmick (2019). « Homogeneous
network embedding for massive graphs via reweighted personalized pagerank. » In: arXiv preprint
arXiv:1906.06826 (cit. on pp. 8, 10).
Yang, Yu, Tian Yu Liu, and Baharan Mirzasoleiman (2022). « Not All Poisons are Created Equal: Robust
Training against Data Poisoning. » In: International Conference on Machine Learning. PMLR, pp. 25154–
25165 (cit. on p. 68).
Yang, Zhilin, William Cohen, and Ruslan Salakhudinov (2016). « Revisiting semi-supervised learning with
graph embeddings. » In: International Conference on Machine Learning. PMLR, pp. 40–48 (cit. on pp. 62,
162, 170).
BIBLIOGRAPHY 115

Yao, Liuquan and Songhao Liu (2024). « New Upper bounds for KL-divergence Based on Integral Norms. »
In: arXiv preprint arXiv:2409.00934 (cit. on p. 177).
Yin, Dong, Ramchandran Kannan, and Peter Bartlett (2019). « Rademacher complexity for adversarially
robust generalization. » In: International conference on machine learning. PMLR, pp. 7085–7094 (cit. on
pp. 19, 81, 84).
Ying, Rex, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L Hamilton, and Jure Leskovec (2018).
« Graph convolutional neural networks for web-scale recommender systems. » In: Proceedings of the
24th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 974–983 (cit. on
p. 3).
Yoo, Jaemin, Sooyeon Shim, and U Kang (2022). « Model-agnostic augmentation for accurate graph
classification. » In: Proceedings of the ACM Web Conference 2022, pp. 1281–1291 (cit. on pp. 183, 186).
You, Yuning, Tianlong Chen, Yongduo Sui, Ting Chen, Zhangyang Wang, and Yang Shen (2020). « Graph
contrastive learning with augmentations. » In: Advances in neural information processing systems 33,
pp. 5812–5823 (cit. on pp. 82, 186).
Yu, Zhaoning and Hongyang Gao (2022). « Molecular representation learning via heterogeneous motif
graph neural networks. » In: International Conference on Machine Learning. PMLR, pp. 25581–25594
(cit. on p. 8).
Zügner, Daniel and Stephan Günnemann (2019). « Certifiable Robustness and Robust Training for Graph
Convolutional Networks. » In: Proceedings of the 25th ACM SIGKDD International Conference on
Knowledge Discovery & Data Mining. ACM (cit. on pp. 56, 57).
Zeng, Hanqing, Muhan Zhang, Yinglong Xia, Ajitesh Srivastava, Rajgopal Kannan, Viktor Prasanna, Long
Jin, Andrey Malevich, and Ren Chen (2020). « Deep graph neural networks with shallow subgraph
samplers. » In: (cit. on p. 41).
Zeng, Xiangxiang, Xinqi Tu, Yuansheng Liu, Xiangzheng Fu, and Yansen Su (2022). « Toward better drug
discovery with knowledge graph. » In: Current opinion in structural biology 72, pp. 114–126 (cit. on p. 81).
Zeng, Zhichen, Ruizhong Qiu, Zhe Xu, Zhining Liu, Yuchen Yan, Tianxin Wei, Lei Ying, Jingrui He, and
Hanghang Tong (2024). « Graph Mixup on Approximate Gromov–Wasserstein Geodesics. » In: Forty-first
International Conference on Machine Learning (cit. on pp. 83, 90, 186).
Zhan, Haoxi and Xiaobing Pei (2021). « Black-box Gradient Attack on Graph Neural Networks: Deeper
Insights in Graph-based Attack and Defense. » In: arXiv preprint arXiv:2104.15061 (cit. on p. 56).
Zhang, Junlong and Yu Luo (2017). « Degree centrality, betweenness centrality, and closeness centrality in
social network. » In: 2017 2nd international conference on modelling, simulation and applied mathematics
(MSAM2017). Atlantis press, pp. 300–303 (cit. on p. 30).
Zhang, Wentao, Zeang Sheng, Yuezihan Jiang, Yikuan Xia, Jun Gao, Zhi Yang, and Bin Cui (2021a).
« Evaluating deep graph neural networks. » In: arXiv preprint arXiv:2108.00955 (cit. on p. 41).
Zhang, Xiang and Marinka Zitnik (2020). « Gnnguard: Defending graph neural networks against adversarial
attacks. » In: Advances in Neural Information Processing Systems 33, pp. 9263–9275 (cit. on pp. 55, 56,
64, 67, 68, 76).
Zhang, Yi, Miaomiao Li, Siwei Wang, Sisi Dai, Lei Luo, En Zhu, Huiying Xu, Xinzhong Zhu, Chaoyun
Yao, and Haoran Zhou (2021b). « Gaussian mixture model clustering with incomplete data. » In: ACM
Transactions on Multimedia Computing, Communications, and Applications (TOMM) 17.1s, pp. 1–14
(cit. on p. 186).
Zhao, Lingxiao and Leman Akoglu (2020). « PairNorm: Tackling Oversmoothing in GNNs. » In: International
Conference on Learning Representations. URL: [Link] (cit. on
p. 41).
116 BIBLIOGRAPHY

Zhao, Weichen, Chenguang Wang, Xinyan Wang, Congying Han, Tiande Guo, and Tianshu Yu (2024).
« Understanding Oversmoothing in Diffusion-Based GNNs From the Perspective of Operator Semigroup
Theory. » In: arXiv preprint arXiv:2402.15326 (cit. on p. 122).
Zhao, Xin, Zeru Zhang, Zijie Zhang, Lingfei Wu, Jiayin Jin, Yang Zhou, Ruoming Jin, Dejing Dou, and Da
Yan (2021). « Expressive 1-lipschitz neural networks for robust multiple graph learning against adversarial
attacks. » In: International Conference on Machine Learning. PMLR, pp. 12719–12735 (cit. on p. 154).
Zhou, Kaixiong, Xiao Huang, Daochen Zha, Rui Chen, Li Li, Soo-Hyun Choi, and Xia Hu (2021). « Dirichlet
energy constrained learning for deep graph neural networks. » In: Advances in Neural Information
Processing Systems 34, pp. 21834–21846 (cit. on p. 19).
Zhou, Wangchunshu, Canwen Xu, Tao Ge, Julian J. McAuley, Ke Xu, and Furu Wei (2020). « BERT Loses
Patience: Fast and Robust Inference with Early Exit. » In: CoRR abs/2006.04152. arXiv: 2006.04152.
URL : [Link] (cit. on p. 42).
Zhu, Dingyuan, Ziwei Zhang, Peng Cui, and Wenwu Zhu (2019). « Robust graph convolutional networks
against adversarial attacks. » In: Proceedings of the 25th ACM SIGKDD international conference on
knowledge discovery & data mining, pp. 1399–1407 (cit. on pp. 57, 63, 64, 68, 69, 76).
Zhu, Jerry, Bryan Gibson, and Timothy T Rogers (2009). « Human rademacher complexity. » In: Advances
in neural information processing systems 22 (cit. on p. 84).
Zhu, Jiong, Yujun Yan, Lingxiao Zhao, Mark Heimann, Leman Akoglu, and Danai Koutra (2020). « Beyond
homophily in graph neural networks: Current limitations and effective designs. » In: Advances in neural
information processing systems 33, pp. 7793–7804 (cit. on pp. 38, 124).
Zou, Xu, Qinkai Zheng, Yuxiao Dong, Xinyu Guan, Evgeny Kharlamov, Jialiang Lu, and Jie Tang (2021).
« TDGIA: Effective Injection Attacks on Graph Neural Networks. » In: Proceedings of the 27th ACM
SIGKDD Conference on Knowledge Discovery & Data Mining. KDD ’21 (cit. on p. 56).
Zügner, Daniel, Amir Akbarnejad, and Stephan Günnemann (2018). « Adversarial attacks on neural
networks for graph data. » In: Proceedings of the 24th ACM SIGKDD international conference on
knowledge discovery & data mining, pp. 2847–2856 (cit. on pp. 19, 55, 56, 62, 68, 77).
Zügner, Daniel and Stephan Günnemann (2019). « Adversarial attacks on graph neural networks via meta
learning. » In: 7th International Conference on Learning Representations (cit. on pp. 56, 61, 64, 67, 68,
76, 155).
Part VI

APPENDICES
A P P E N D I X : R E T H I N K I N G G R A P H S H I F T O P E R AT O R S I N G N N S F O R G R A P H
a
R E P R E S E N TAT I O N L E A R N I N G

119
120 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

A .1 D ATA S E T S A N D I M P L E M E N TAT I O N D E TA I L S

In this section, we present the benchmark datasets used for our experiments, and the process used for
the training.

A .1.1 Statistics of the Node Classification Datasets

We use ten widely used datasets in the GNN literature. In particular, we run experiments on the node
classification task using the citation networks Cora, CiteSeer, and PubMed (Sen et al., 2008), the co-
authorship networks CS and Physiscs (Shchur et al., 2018b), the citation network between Computer
Science arXiv papers OGBN-Arxiv (Hu et al., 2020a), the Amazon Computers and Amazon Photo networks
(Shchur et al., 2018b), the non-homophilous datasets Penn94 (Traud, Mucha, and Porter, 2012), genius
(Lim and Benson, 2021), deezer-europe (Rozemberczki and Sarkar, 2020) and arxiv-year (Hu et al., 2020a),
and the disassortative datasets Chameleon, Squirrel (Rozemberczki, Allen, and Sarkar, 2021), and Cornell,
Texas, Wisconsin from the WebKB dataset (Lim et al., 2021a). Characteristics and information about the
datasets utilized in the node classification part of the study are presented in Table a.1.

Table a.1 – Statistics of the node classification datasets used in our experiments.

Dataset #Features #Nodes #Edges #Classes Edge Homophily


Cora 1,433 2,708 5,208 7 0.809
CiteSeer 3,703 3,327 4,552 6 0.735
PubMed 500 19,717 44,338 3 0.802
CS 6,805 18,333 81,894 15 0.808
arxiv-year 128 169,343 1,157,799 5 0.218
chameleon 2,325 2,277 62,792 5 0.231
Cornell 1,703 183 557 5 0.132
deezer-europe 31,241 28,281 185,504 2 0.525
squirrel 2,089 5,201 396,846 5 0.222
Wisconsin 1,703 251 916 5 0.206
Texas 1,703 183 574 5 0.111
Photo 745 7,650 238,162 8 0.827
ogbn-arxiv 128 169,343 2,315,598 40 0.654
Computers 767 13752 491,722 10 0.777
Physics 8,415 34,493 495,924 5 0.931
Penn94 4,814 41,554 2,724,458 3 0.470

A .1.2 Implementation Details

We train all the models using the Adam optimizer (Kingma and Ba, 2015b). To account for the impact of
random initialization, each experiment was repeated 10 times, and the mean and standard deviation of the
A .1 D ATA S E T S A N D I M P L E M E N TAT I O N D E TA I L S 121

results were reported. The experiments have been run on both a NVIDIA A100 GPU and a RTX A6000
GPU.
Training of our CGNN. We train our model using the Adam optimizer (Kingma and Ba, 2015b), with a
weight decay on the parameters of 5 × 10−4 , an initial learning rate of 0.005 for the exponential parameters
and an initial learning rate of 0.01 for all other model parameters. We repeated the training 10 times to
test the stability of the model. We tested 7 initialization of the weights (m1 , m2 , m3 , e1 , e2 , e3 , a). These
initializations are reported in Table a.2 in Appendix a.1, and correspond to classical GSOs when the
chosen centrality is the degree. For the Cora, CiteSeer, and Pubmed datasets, we used the provided
train/validation/test splits. For the remaining datasets, we followed the framework of Lim et al. (2021a) and
Rozemberczki, Allen, and Sarkar (2021).

A .1.3 Weights Initialization

In this part, we present the different initializations of CGSO. When the chosen centrality is the degree, i.e.
V = D, the initializations corresponds to popular classical GSO (Dasoulas, Lutzeyer, and Vazirgiannis,
2021b).

Table a.2 – Differenet initialization of the weights (m1 , m2 , m3 , e1 , e2 , e3 , a).

Initialization of (m1 , m2 , m3 , e1 , e2 , e3 , a) Corresponding GSO Description when V = D

(0, 1, 0, 0, 0, 0, 0) A(V) = A Adjacency matrix


(1, −1, 0, 1, 0, 0, 0) L(V) = V − A Unnormalised Laplacian matrix
(1, 1, 0, 1, 0, 0, 0) Q(V) = V + A Signless Laplacian matrix
(0, −1, 1, 0, −1, 0, 0) Lrw (V) = I − V −1 A Random-walk Normalised Laplacian
(0, −1, 1, 0, −1/2, −1/2, 0) Lsym (V) = I − V−1/2 AV −1/2 Symmetric Normalised Laplacian
(0, 1, 0, 0, −1/2, −1/2, 1) Â(V) = V−1/2 A 1 V−1/2 Normalised Adjacency matrix
(0, 1, 0, 0, −1, 0, 0) H(V) = V −1 A Mean Aggregation Operator

A .1.4 Hyperparameter Configurations

For a more balanced comparison, however, we use the same training procedure for all the models. The
hyperparameters in each dataset where performed using a Grid search on the classical GCN (i.e. with the
GSO : Normalised adjacency) over the following search space:
— Hidden size : [16, 32, 64, 128, 256, 512],
— Learning rate : [0.1, 0.01, 0.001],
— Dropout probability: [0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8].
The number of layers was fixed to 2. The optimal hyperparameters can be found in Table a.3.
122 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

Table a.3 – Hyperparameters used in our experiments.

Dataset Hidden Size Learning Rate Dropout Probability

Cora 64 0.01 0.8


CiteSeer 64 0.01 0.4
PubMed 64 0.01 0.2
CS 512 0.01 0.4
arxiv-year 512 0.01 0.2
chameleon 512 0.01 0.2
cornell 512 0.01 0.2
deezer-europe 512 0.01 0.2
squirrel 512 0.01 0.2
Wisconsin 512 0.01 0.2
Texas 512 0.01 0.2
Photo 512 0.01 0.6
OGBN-Arxiv 512 0.01 0.5
Computers 512 0.01 0.2
Physics 512 0.01 0.4
Penn94 64 0.01 0.2

A .2 A D D I T I O N A L R E S U LT S F O R T H E N O D E C L A S S I F I C AT I O N TA S K

To further evaluate our CGCN and CGATv2, we compute its performance on additional datasets. The
results of this study are presented in Table a.4.

A .3 S I M P L E G R A P H C O N VO L U T I O N A L N E T WO R K S

In Tables a.5 and a.6, we present the results of our centrality-aware Simple Graph Convolutional Networks
CGSC of 2 layers. As noticed in most cases, by incorporating our CGSO, we outperform the classical SGC.
To also understand the effect of the centrality on the oversmoothing effect, we analyzed the variation of
Dirichlet Energy (Zhao et al., 2024) of CGSC across different numbers of layers. As noticed, while the
centrality has a lower effect on the oversmoothing in the homophilous dataset Cora, we notice a larger
impact on the heterophilious dataset Chameleon.

A .4 COMBINING LOCAL AND GLOBAL CENTRALITIES

A .5 T H E G R A P H S T RU C T U R E O F T H E S TO C H A S T I C B L O C K B A R A B Á S I – A L B E RT M O D E L S

In Figure a.2, we give an example of an adjacency matrix sampled from SBBAM model, presented
previously in Section 3.4.1. Yellow points indicate edges, while purple points represent non-edges. Notably,
A .6 k- C O R E D I S T R I BU T I O N I N S TO C H A S T I C B L O C K B A R A B Á S I – A L B E RT M O D E L S 123

Table a.4 – Classification accuracy (± standard deviation) of the models on different benchmark node classification
datasets. The higher the accuracy (in %) the better the model. ⃝1 GCN Based models ⃝ 2 Other Vanilla
GNN baselines ⃝ 3 CGCN ⃝ 4 CGATv2. Highlighted are the first, second best results. OOM means Out of
memory.
Model Cora Texas Photo ogbn-ariv CS Computers Physics Penn94
GCN w/ A 78.61±0.51 63.51±2.18 82.31±2.61 13.23±6.44 87.70±1.25 69.32±3.64 88.92±1.93 52.35±0.36
GCN w/ L 31.57±0.41 84.32±2.65 27.42±6.23 10.91±1.49 23.75±3.22 26.27±3.89 35.31±3.71 65.31±0.59
GCN w/ Q 77.32±0.50 60.54±1.32 77.06±6.73 10.50±1.97 89.42±1.31 47.72±18.37 90.69±2.13 53.46±2.16

1 GCN w/ Lrw 26.59±1.11 78.38±2.09 24.60±4.21 8.07±0.07 26.34±4.09 13.76±3.96 28.19±3.75 69.82±0.44
GCN w/ Lsym 26.79±0.50 71.35±1.32 22.82±2.67 20.18±0.24 24.39±1.96 16.06±5.19 30.94±3.11 70.57±0.30
GCN w/ Â 80.84±0.40 60.81±1.81 78.94±1.65 65.80±0.14 91.52±0.75 68.91±3.00 93.72±0.80 74.60±0.42
GCN w/ H 80.15±0.37 59.46±0.00 73.95±4.75 63.34±0.15 90.98±1.84 62.01±4.36 92.16±1.12 71.78±0.47

GIN 79.06±0.47 57.03±1.89 83.00±2.52 9.30±6.42 89.53±1.20 55.89±13.45 89.15±2.44 OOM


GAT 77.73±1.83 52.16±6.74 71.56±3.48 67.36±0.13 67.67±3.96 59.73±3.59 80.91±4.48 73.85±1.38

2 GATv2 74.53±2.48 48.11±3.78 73.49±2.49 68.14±0.07 70.13±4.92 58.18±4.76 83.28±3.68 75.54±2.54
PNA 56.67±10.53 63.51±4.05 16.75±5.59 OOM OOM 13.62±6.39 OOM OOM
CGCN w/ D 79.45±0.58 81.89±9.38 88.78±1.74 69.09±0.21 91.28±1.29 79.26±1.87 92.51±1.16 73.06±0.34
CGCN w/ Vcore 79.80±0.43 77.84±5.51 88.53±1.40 65.54±0.57 91.37±1.18 77.35±2.67 91.98±1.49 78.11±3.74

3
CGCN w/ Vℓ-walks 79.52±0.35 78.11±5.82 83.72±2.03 22.54±8.22 89.87±1.20 68.56±3.39 89.84±2.74 68.44±0.37
CGCN w/ VPR 79.51±15.01 82.70±4.95 81.28±6.08 68.56±0.18 88.76±30.68 65.54±6.43 89.64±10.3 72.59±0.84

CGATv2 w/ D 79.07±0.64 82.70±5.30 87.97±1.77 70.09±0.10 91.48±1.05 78.62±2.35 91.32±1.18 72.81±0.36


CGATv2 w/ Vcore 79.03±0.96 83.78±6.62 89.72±1.54 69.93±0.13 91.91±1.06 77.31±3.33 91.15±1.07 72.86±0.41

4
CGATv2 w/ Vℓ-walks 78.58±0.58 79.73±4.72 88.11±2.02 70.51±0.24 90.73±1.46 79.09±1.66 89.98±1.37 72.79±0.43
CGATv2 w/ VPR 78.60±0.38 83.78±4.98 88.38±2.09 69.26±0.12 91.77±1.00 74.95±3.05 92.73±1.44 75.16±0.69

there are variations in edge density across different blocks: while the first block appears to be predominantly
heterophilic, the third block inherits a homophilic cluster structure.

A .6 k- C O R E D I S T R I BU T I O N I N S TO C H A S T I C B L O C K B A R A B Á S I – A L B E RT M O D E L S

In Figure a.3, we illustrate the k-core distribution of the three individual blocks, each corresponding to BA
graphs, and the combined SBBAM graph.

A .7 SPECTRAL CLUSTERING ALGORITHM

In Algorithm 5, we outline the steps of the Spectral Clustering Algorithm using CGSOs. This algorithm is
applied on SBBAMs for cluster recovery and on the Cora dataset for centrality recovery.

A .8 A D D I T I O N A L R E S U LT S F O R T H E S P E C T R A L C L U S T E R I N G TA S K

In this section, we report the ARI value of the spectral clustering task described in Section 3.4.
124 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

Table a.5 – Classification accuracy (± standard deviation) of the models on different benchmark node classification
datasets. The higher the accuracy (in %) the better the model. ⃝1 CSGC with nodes centrality, ⃝
2 SGC.
Highlighted are the best results.
Model CiteSeer PubMed arxiv-year chamelon Cornell deezer-europe squirrel Wisconsin
CSGC w/ D 67.70±0.17 77.37±0.25 35.01±0.16 59.10±1.66 72.70±4.26 58.22±0.47 40.12±1.69 75.88±4.96
CSGC w/ Vcore 66.85±0.15 78.19±0.12 37.71±0.17 63.11±4.56 72.16±5.55 61.29±0.50 38.66±2.27 75.10±4.12

1
CSGC w/ Vℓ-walks 67.09±0.05 77.50±0.18 36.67±0.22 45.26±2.51 74.32±6.07 59.69±0.50 27.85±1.38 81.76±3.73
CSGC w/ VPR 64.91±0.47 76.47±0.37 23.87±0.51 55.18±3.36 69.46±5.14 58.94±0.48 26.73±2.25 75.29±4.31

2 SGC 64.96±0.10 75.72±0.12 26.61±0.24 38.44±4.41 45.41±5.77 62.66±0.48 19.88±0.79 53.53±8.09

Table a.6 – Classification accuracy (± standard deviation) of the models on different benchmark node classification
datasets. The higher the accuracy (in %) the better the model. ⃝1 CSGC with nodes centrality, ⃝
2 SGC.
Highlighted are the best results.
Model Cora Texas Photo ogbn-arxiv CS Computers Physics Penn94
CSGC w/ D 80.10±0.11 76.76±3.24 89.38±1.81 67.94±0.06 92.29±1.04 79.04±1.94 92.32±1.2 78.84±4.15
CSGC w/ Vcore 78.80±0.17 77.30±3.86 88.58±1.68 62.54±0.16 91.82±1.10 76.46±2.29 91.71±1.63 76.25±1.21

1
CSGC w/ Vℓ-walks 77.32±0.29 80.27±5.41 88.78±2.69 66.41±0.05 91.96±0.84 76.17±4.92 91.71±1.58 73.20±0.36
CSGC w/ VPR 76.92±0.39 77.30±4.86 84.33±3.06 44.82±1.16 90.24±0.86 61.51±2.71 91.57±1.70 77.24±0.67

2 SGC 78.79±0.13 58.65±4.20 24.0±11.82 60.48±0.14 70.78±5.47 11.34±11.67 91.69±1.48 66.63±0.62

A .9 C G N N W I T H H E T E R O P H I LY

In this section, we incorporate our learnable CGSOs into H2GCN Zhu et al., 2020, designed for
heterophilic graphs. We compared the results of CH2GCN and H2GCN on datasets with low homophily.
We report the results of this experiment in Table a.8. As noticed, our CH2GCN outperforms H2GCN.

A .10 PROOFS OF PROPOSITIONS

In this section, we details the proofs of the propositions 3.3.1, 3.3.2 and 3.3.4.
A .10 P R O O F S O F P R O P O S I T I O N S 125

(b) Cora (b) Chameleon


Degree Degree
107 k core k core
count. walks
105 count. walks
Dirichlet Energy

Dirichlet Energy
106
102
105
10 1

104 10 4

0 2 4 6 8 10 0 2 4 6 8 10
Layers Layers
Figure a.1 – Dirichlet Energy variation with layers in (a) Cora and (b) Chamelon.

Table a.7 – Classification accuracy (± standard deviation) of the models on different benchmark node classification
datasets. The higher the accuracy (in %) the better the model.
Model Cora Texas Photo ogbn-arxiv CS Computers Physics Penn94
CGCN w/ D 79.45±0.58 81.89±9.38 88.78±1.74 69.09±0.21 91.28±1.29 79.26±1.87 92.51±1.16 73.06±0.34
CGCN w/ Vcore 79.80±0.43 77.84±5.51 88.53±1.40 65.54±0.57 91.37±1.18 77.35±2.67 91.98±1.49 78.11±3.74
CGCN w/ Vℓ-walks 79.52±0.35 78.11±5.82 83.72±2.03 22.54±8.22 89.87±1.20 68.56±3.39 89.84±2.74 68.44±0.37
CGCN w/ VPR 79.51±15.01 82.70±4.95 81.28±6.08 68.56±0.18 88.76±30.68 65.54±6.43 89.64±10.3 72.59±0.84

CGCN w/ D & Vcore 79.88±0.38 78.92±4.32 89.06±1.28 67.67±0.26 91.63±0.95 78.41±1.94 91.28±3.17 80.28±2.93
CGCN w/ D & Vℓ-walks 79.38±0.72 81.89±4.69 86.78±2.75 69.57±0.24 91.78±1.04 78.39±2.36 91.2±1.56 72.5±0.48
CGCN w/ D & VPR 79.84±0.4 78.11±2.55 82.76±2.06 21.28±9.89 90.04±0.57 65.66±4.96 90.24±1.86 71.03±5.83

A .10.1 Proof of Proposition 3.3.1

Proof of Proposition 3.3.1. We first prove that the operator MG is self-adjoint.


For φ1 , φ2 ∈ L2 ( G ), we have:

< M G φ1 , φ2 > G = ∑ v(i) (MG φ1 ) (i) φ¯2 (i)


i ∈V
!
1
= ∑ v (i ) v (i ) ∑ φ1 ( j ) φ¯2 (i )
i ∈V j∈Ni
!
= ∑ φ¯2 (i) ∑ φ1 ( j )
i ∈V j∈Ni

= ∑∑ ai,j φ¯2 (i ) φ1 ( j)
i ∈V j∈Ni

= ∑ ai,j φ¯2 (i ) φ1 ( j).


i,j∈V
126 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

0
50
100
150
200
250

0 50 100 150 200 250


Figure a.2 – The adjacency matrix of the synthetic graph generated from an SBBAM with 3 blocks.

Algorithm 5: Spectral Clustering using the Centrality GSOs


1 Inputs: Graph G, Centrality GSO Φ, Number of clusters to retrieve C.
1. Compute the eigenvalues {λ}in=1 and eigenvectors {u}in=1 of Φ;
2. Consider only the eigenvectors U ∈ R N ×C corresponding to the C largest eigenvalues;
3. Cluster rows of U, corresponding to nodes in the graph, using the K-Means algorithm to retrieve a node
partition P with C clusters;
P = K-Means(U, C )
return P ;

Similarly, we also have that,

< φ1 , M G φ2 > G = ∑ v(i) φ1 (i)(MG φ2 )(i)


i ∈V
!
1
= ∑ v ( i ) φ1 ( i ) ∑ φ2 ( j )
i ∈V
v ( i) j∈Ni
!
= ∑ φ1 ( i ) ∑ φ2 ( j )
i ∈V j∈Ni

= ∑ ai,j φ¯2 (i ) φ1 ( j)
i,j∈V

= ∑ a j,i φ¯2 ( j) φ1 (i ).
i,j∈V

Thus,
A .10 P R O O F S O F P R O P O S I T I O N S 127

(a) Seperate BA Graphs (b) Combined Graph


1.0 q=5
q = 10
0.6
q = 15

0.8 0.5
0.4
0.6
Density

Density
0.3
0.4
0.2
0.2 0.1
0.0 2 4 6 8 10 12 14 0.0 2 4 6 8 10 12 14
KCore KCore

Figure a.3 – The left figure represents the k-core distributions of three different BA models with the hyperparameters
r = 5, 10 and 15 serving as blocks of of our SBBAM. The right figure represents the k-core distribution of
the SBBAM.

1.5
Degree 1.5
K-Core 1.5
Count of Walks 1.5
PageRank
2.5 1.35
2.0 4
1.0 1.0 1.0 1.0
2.0
1.5 1.30
3
0.5 0.5 0.5 1.5 0.5
1.0
2 1.0 1.25
0.5
e2

e2

e2

e2
0.0 0.0 0.0 0.0
1 0.5
0.0 1.20
0.5 0.5 0.5 0.0 0.5
0.5 0
1.0 1.0 1.0 0.5 1.0 1.15
1.0 1
1.0
1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 1.5 1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 1.51.5 1.0 0.5 0.0 0.5 1.0 1.5 1.10
e3 e3 e3 e3
Figure a.4 – Result for the spectral clustering task on Cora graph with core numbers considered as clusters. We
report the values of the Adjusted Rand Information (ARI) in % different combination of the exponents
(e2 , e3 ) in Ve2 AVe3 .

(
2 < MG φ1 , φ2 >G = ∑i,j∈V ai,j φ¯2 (i ) φ1 ( j),
∀ φ1 , φ2 ∈ L ( G ) , (a.1)
< φ1 , MG φ2 >G = ∑i,j∈V a j,i φ¯2 ( j) φ1 (i ).
Since ai,j = a j,i , we conclude that MG is self-adjoint, i.e.

< MG φ1 , φ2 >G =< φ1 , MG φ2 >G .


MG is self-adjoint, the space L2 ( G ) is finite-dimensional, thus is diagonalizable in an orthonormal basis,
and its eigenvalues are real.
We define the following norm,
⟨MG φ, φ⟩G
∥MG ∥ = sup .
φ ̸ =0 ∥ φ ∥2
We will now prove that all eigenvalues have absolute values at most γ = mini∈V v(i )/deg(i ). For that, we
will first compute the two inner-products < ( I − MG ) φ, φ >G and < ( I + MG ) φ, φ >G . For any φ ∈ L2 ( G ),
using (a.1), we have that:
(
< φ, φ >G = ∑i∈V v(i )| φ(i )|2 ,
< MG φ, φ >G = ∑i,j∈V a(i, j) φ̄(i ) φ( j).
 
v (i )
Let’s first take the simple case, where γ = mini∈V deg(i) ≤ 1, then,
128 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

Table a.8 – Classification accuracy (± standard deviation) of the models on different benchmark node classification
datasets. The higher the accuracy (in %) the better the model. ⃝ 1 CH2GCN with nodes centrality,

2 H2GCN. Highlighted are the best results.

Model Texas Cornell Wisconsin chameleon


CH2GCN w/ D 79.73±5.02 68.65±5.16 79.80±4.02 67.89±4.23
CH2GCN w/ Vcore 78.92±5.77 68.92±7.28 79.80±3.40 60.00±5.63

1
CH2GCN w/ Vℓ-walks 78.11±6.10 68.92±6.19 82.35±5.04 44.28±2.32
CH2GCN w/ VPR 60.27±5.41 44.86±7.76 52.35±7.75 31.95±5.79

2 H2GCN 56.76±6.73 51.08±6.89 55.29±5.10 63.93±2.07

2 < φ, φ >G = 2 ∑ v(i )| φ(i )|2


i ∈V
≥ 2γ ∑ deg(i )| φ(i )|2
i ∈V
≥ 2 ∑ deg(i )| φ(i )|2
i ∈V
!
≥2∑ ∑ a(i, j) | φ(i )|2
i ∈V j∈V

≥2 ∑ a(i, j)| φ(i )|2 .


i,j∈V

Therefore,
2 < (I − MG ) φ, φ >G = 2 < φ, φ >G −2 < MG φ, φ >G
≥2 ∑ a(i, j)| φ(i )|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)
i,j∈V i,j∈V

≥ ∑ a(i, j)| φ(i )|2 + ∑ a(i, j)| φ( j)|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)


i,j∈V i,j∈V i,j∈V

≥ ∑ a(i, j)| φ(i ) − φ( j)|2 .


i,j∈V

Similarly, we can prove that,

2 < (I + MG ) φ, φ >G ≥ ∑ a(i, j)| φ(i ) + φ( j)|2 .


i,j∈V

Therefore, if ϕ ̸= 0, then,
(
< (I − MG ) φ, φ >G ≥ 0,
⇒< φ, φ >G ≤< MG φ, φ >G ≤< φ, φ >G .
< (I + MG ) φ, φ >G ≥ 0
| < MG φ, φ >G |
⇒ ≤ 1.
< φ, φ >G
A .10 P R O O F S O F P R O P O S I T I O N S 129

Thus, ∥MG ∥ ≤ 1, i.e. all the eigenvalues have absolute values at most 1. Let now consider the general
e = 1 V = diag( v(1) , . . . , v( N ) ). Since,
case, where γ is not necessarily smaller than 1. Let’s consider V γ γ γ

!
v(˜i )
γ̃ = min
i ∈V deg(i )
 
1 v (i )
=
γ deg(i )
γ
=
γ
= 1.
e −1 A = 1 V−1 A = 1 MG have absolute values at most 1. Thus, all
MG = V
Therefore, all the eigenvalues of f γ γ
the eigenvalues of MG have absolute values at most γ.

A .10.2 Proof of Proposition 3.3.2

Proof of Proposition 3.3.2. We will prove the first property.


We consider P as the number of connected components, i.e. G = iP=1 Ci .
S

The adjacency matrix of the graph G is,


 
A C1 0 0
..
 
.
 
 
0 0
 
A=  A Ci .

 .. 

 . 

0 0 ACP

And the transformation of A by the Markov Average operator MG is,


 
M
 C1
0 0
..

.
 
 
0 0
 
MG =   MC i .

 .. 

 . 

0 0 MC P

According to Proposition 3.3.1, for each connected component Ci , the matrix MCi is diagonalizable
in an orthonormal basis, and its eigenvalues are real numbers. We denote by eCi = [eC1 i , . . . , eC|Ci | ] the
i

eigenvectors basis of MCi corresponding the eigenvalues λ Ci = [λ1Ci , . . . , λC|Ci | ].


i
We consider the set of vectors
 
e C1 0 0
..
 
.
 
 
0 0
 
e= e Ci .

 .. 

 . 

0 0 e CP
130 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

The column vectors of e are eigenvectors of the matrix MG , and which achieves the conditions of Property
1. Let’s now prove the formulas of the mean and standard deviation of the MG spectrum. The matrix V−1 A
is defined as follow,
1
∀1 ≤ i, j ≤ N, (D−1 A)i,j = Ai,j .
v (i )
2
Therefore, the diagonal elements of the matrix D−1 A is defined as follow,
 2   
−1
∀1 ≤ i ≤ N, D A = D−1 AD−1 A
i,i i,i

= ∑(D−1 A)i,k (D−1 A)k,i


j
Ai,j A j,i
=∑
j
v (i ) × v ( j )
A2i,j
=∑
j
v (i ) × v ( j )
1
= ∑ v (i ) × v ( j )
.
j∈Ni

Thus,
 h i
µ (MG ) = Mean Spectrum V−1 A
1  h i
= Sum Spectrum V−1 A
N
1 N
N i∑
= (D−1 A)i,i
=1
N
1 1
=
N ∑ v(i) Ai,i
i =1
N
1 1
=
N ∑ v (i ) ,
i =1
A .10 P R O O F S O F P R O P O S I T I O N S 131

and,
 h i
σ (MG ) = Stdev Spectrum V−1 A
v
u1
u

2
=t λ − Mean spϕ
N λ∈Spectrum[V−1 A]
v 
u
u 1
u

2
=t λ2  − Mean spϕ
N λ∈Spectrum[V−1 A]
s 
1 2
2
= Sum(Spectrum [ϕ ]) − Mean spϕ
N
s
1 h i  2
2
= Sum(Spectrum (D−1 A) ) − Mean spϕ
N
s 
1 h −1 2 i 2
= Tr (D A) − Mean spϕ
N
v !
u
u 1 1
∑ ∑
2
= t − Mean spϕ
N i=1 j∈N v(i ) × v( j)
i
v 
u
u 1 1
u
∑  − Mean spϕ 2 .

=
N (i,j)∈E v(i ) × v( j)
t

A .10.3 Proof of Proposition 3.3.4

Proof of Proposition 3.3.4. Let W ⊂ V, such that |W | ≤ 12 |V |.


132 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

|W |
For φ = 1W − µG (W ) where µG (W ) = Nv and Nv = ∑i∈V v(i ) = |V |v .

2 < (I − MG ) φ, φ >G = 2 < φ, φ >G −2 < MG φ, φ >G


≤ 2 ∑ v(i )| φ(i )|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)
i ∈V i,j∈V

≤ 2 ∑ β × deg(i )| φ(i )|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)


i ∈V i,j∈V

≤ 2 ∑ deg(i )| φ(i )|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)


i ∈V i,j∈V
!
≤2∑ ∑ a(i, j) | φ(i )|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)
i ∈V j∈V i,j∈V

≤2 ∑ a(i, j)| φ(i )|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)


i,j∈V i,j∈V

≤ ∑ a(i, j)| φ(i )|2 + ∑ a(i, j)| φ( j)|2 − 2 ∑ a(i, j) φ̄(i ) φ( j)


i,j∈V i,j∈V i,j∈V

≤ ∑ a(i, j)| φ(i ) − φ( j)|2


i,j∈V

≤ ∑ a(i, j)|1W (i ) − 1W ( j)|2 .


i,j∈V

The non-zero terms in ∑i,j∈V a(i, j)| φ(i ) − φ( j)|2 are those where i and j are adjacent, but one of them is
in W and the other not.

1
2 i,j∑
< (I − MG ) φ, φ >G ≤ a(i, j)|1W (i ) − 1W ( j)|2
∈V

= # E (W ) .

1
There 2 was removed because of the symmetry. We also have that,

1 1
Nv
< 1W , 1W >G =
Nv ∑ v ( i ) = µ G (W ) ,
i ∈V

and,
1 1
∑ v(i)µG (W ) = (µG (W ))
2
< 1W , µG (W ) >G = ,
Nv Nv i ∈V
A .10 P R O O F S O F P R O P O S I T I O N S 133

Therefore,
1 1
< φ, φ >G = < 1W − µG (W ), 1W − µG (W ) >G
Nv Nv
1 1
= < 1W , 1W − µG (W ) >G − < µG (W ), 1W − µG (W ) >G
Nv Nv
1 2 1
= < 1W , 1W >G − < 1W , µG (W ) >G + < µ G (W ) , µ G (W ) > G
Nv Nv Nv
1
= µG (W ) − 2 (µG (W ))2 + (µG (W ))2 < 1, 1 >G
Nv
1
Nv i∑
= µG (W ) − 2 (µG (W ))2 + (µG (W ))2 v (i )
∈V
Nv
= µG (W ) − 2 (µG (W ))2 + (µG (W ))2
Nv
= µG (W ) − (µG (W ))2
= µG (W ) (1 − µG (W ))
= µ G (W ) µ G (W ′ ) ,

where W ′ = V − W.
By definition,
< (I − MG ) φ̃, φ̃ >G
λ1 ( G ) = min .
φ̸̃=0 < φ̃, φ̃ >G
Therefore,
< (I − MG ) φ, φ >G
λ1 ( G ) ≤
< φ, φ >G
# E (W ) Nv

Nv < φ, φ >G
# E (W ) Nv
≤ .
Nv µG (W )µG (W ′ )
Since,
v − |W | v v − |W | v
≤ µ G (W ) ≤ ,
v+ |V |v v+ |V |v
then,
|W | v v − |W ′ | v
Nv µG (W )µG (W ′ ) ≥ Nv
|V |v v+ |V |v
∑ v (i ) v − |W ′ | v
≥ i∈V |W | v
|V |v v+ |V |v
∑ v (i ) v − |W ′ | v
≥ i∈V |W | v
∑i∈V v(i ) v+ |V |v

v − |W | v
≥ |W | v
v+ |V |v
v− |W ′ | v
≥ |W | v
v+ |V |v
v−
≥ |W | v ,
2v+
134 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

because (
|W |v ≤ 12 |V |v 1
⇒ |W ′ |v ≥ |V |v .
W′ = V −W 2

Thus,
1 2v+ #E (W )
∀W ⊂ V , |W |v ≤ |V |v ⇒ λ1 ( G ) ≤ .
2 v − |W | v
Thus,
2v+ v2
λ1 ( G ) ≤ Nv h( G ) ≤ 2N + hv ( G ).
v− v−

A .11 AV E R A G E D E G R E E O F A B A R A B A S I – A L B E R T M O D E L

In this section, we details the proofs of the propositions 3.4.1.

Proof of Proposition 3.4.1. We start with a small graph of N0 nodes and r0 edges. At each time step, we
increase the number of edges by r. Thus, if N is the number of nodes at a certain time step, then there are
exactly r0 + r ( N − N0 ) edges.
As each edge contributes to the degree of two nodes, thus, the average degree is twice the number of
edges divided by the number of nodes N. Therefore,

2
deg( G BA ) = (r0 + r ( N − r0 ))
N
r0 r
= 2r + 2 − 2N0 .
N N

A .12 L E A R N E D PA R A M E T E R S O F D I F F E R E N T C E N T R A L I T Y B A S E D G S O S

In this section, we present some graph properties of the used dataset. We specifically present the node
density, the homophily coefficient as well as the average value of different centrality metrics in Table a.9.
We also present the (m1 , m2 , m3 , e1 , e2 , e3 , a) learned by the GNN in Tables a.10, a.11, a.12 and a.13.
A .12 L E A R N E D PA R A M E T E R S O F D I F F E R E N T C E N T R A L I T Y B A S E D G S O S 135

Table a.9 – Detailed graph properties of the used datasets.

Dataset density Avg. Degree Avg. PageRank Avg. K-core Avg. Count. Paths homophily

Physics 4.16 × 10−4 14.37 2.89 × 10−5 7.71 449.22 0.931


Photo 4.07 × 10−3 31.13 1.30 × 10−4 16.97 3204.098 0.827
Cora 1.43 × 10−3 3.89 3.69 × 10−4 2.31 42.52 0.809
CS 4.87 × 10−4 8.93 5.45 × 10−5 4.94 162.75 0.808
PubMed 2.28 × 10−4 4.49 5.07 × 10−5 2.39 75.43 0.802
Computers 2.60 × 10−3 35.75 7.27 × 10−5 18.84 6221.39 0.777
CiteSeer 8.22 × 10−4 2.73 3.00 × 10−4 1.73 18.91 0.735
ogbn-arxiv 8.07 × 10−5 13.67 5.90 × 10−6 7.13 4898.16 0.654
deezer-europe 2.31 × 10−4 6.55 3.53 × 10−5 3.57 106.16 0.525
Penn94 1.57 × 10−3 65.56 2.40 × 10−5 33.68 10662.08 0.470
chameleon 1.21 × 10−2 27.57 4.39 × 10−4 16.60 2913.48 0.231
squirrel 1.46 × 10−2 76.30 1.92 × 10−4 41.55 31888.02 0.222
arxiv-year 8.07 × 10−5 6.88 5.90 × 10−6 7.13 82.85 0.218
Wisconsin 1.48 × 10−2 3.64 3.98 × 10−3 2.05 76.26 0.206
Cornell 1.68 × 10−2 3.04 5.46 × 10−3 1.74 58.47 0.132
Texas 1.77 × 10−2 3.13 5.46 × 10−3 1.71 70.72 0.111

A .12.1 Degree Centrality

A .12.2 k-Core Centrality

A .12.3 PageRank Centrality


136 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

Table a.10 – Graph Properties of the used datasets and the corresponding learned hyperparameters in GAGCN w/
Degree
Dataset Graph Properties Hyperparameters

Avg. K-core Avg. Count. Walks homophily e1 e2 e3 m1 m2 m3 a

Physics 7.71 449.22 0.931 0.28 (0.01) −0.31 (0.00) −0.32 (0.00) 0.34 (0.01) 1.33 (0.01) 0.31 (0.01) 1.36 (0.01)
Photo 16.97 3204.098 0.827 0.39 (0.06) −0.26 (0.01) −0.25 (0.01) 0.59 (0.05) 1.51 (0.01) 0.53 (0.04) 1.70 (0.04)
Cora 2.31 42.52 0.809 0.31 (0.04) 0.02 (0.01) −0.02 (0.01) 0.67 (0.02) 1.43 (0.04) 0.66 (0.02) 0.69 (0.01)
CS 4.94 162.75 0.808 0.33 (0.00) −0.25 (0.00) −0.26 (0.00) 0.44 (0.01) 1.44 (0.00) 0.40 (0.01) 1.47 (0.01)
PubMed 2.39 75.43 0.802 0.28 (0.01) −0.27 (0.00) −0.28 (0.00) 0.39 (0.00) 1.40 (0.01) 0.38 (0.00) 1.39 (0.01)
Computers 18.84 6221.39 0.777 0.40 (0.05) −0.74 (0.02) 0.24 (0.03) 0.74 (0.05) 1.60 (0.05) 0.66 (0.04) 0.86 (0.10)
CiteSeer 1.73 18.91 0.735 0.35 (0.00) −0.21 (0.01) −0.22 (0.01) 0.49 (0.01) 1.49 (0.01) 0.47 (0.01) 1.50 (0.01)
ogbn-arxiv 7.13 4898.16 0.654 −0.08 (0.02) −0.29 (0.01) −0.41 (0.00) 0.13 (0.01) 1.31 (0.04) 0.13 (0.01) 1.00 (0.01)
deezer-europe 3.57 106.16 0.525 0.31 (0.04) −0.51 (0.03) −0.54 (0.02) 0.59 (0.04) −0.96 (0.04) 1.55 (0.03) −0.59 (0.03)
Penn94 33.68 10662.08 0.470 0.51 (0.01) −1.00 (0.02) −0.09 (0.02) 0.98 (0.01) 1.01 (0.04) 0.82 (0.01) 0.95 (0.03)
chameleon 16.60 2913.48 0.231 0.15 (0.04) −0.06 (0.01) −0.06 (0.01) −0.17 (0.03) 0.88 (0.02) −0.16 (0.03) −0.15 (0.02)
squirrel 41.55 31888.02 0.222 0.38 (0.05) −0.26 (0.03) −0.24 (0.03) 0.31 (0.80) 1.75 (0.07) 0.26 (0.70) 1.69 (0.56)
arxiv-year 7.13 82.85 0.218 −0.25 (0.01) −0.27 (0.01) −0.40 (0.01) 0.01 (0.01) 0.99 (0.01) 0.05 (0.01) 0.80 (0.01)
Wisconsin 2.05 76.26 0.206 0.95 (0.05) −0.09 (0.04) −0.05 (0.01) 1.27 (0.25) −0.94 (0.05) 0.66 (0.07) −0.64 (0.06)
Cornell 1.74 58.47 0.132 0.88 (0.05) −0.17 (0.07) −0.07 (0.03) 1.04 (0.29) −0.86 (0.11) 0.80 (0.10) −0.78 (0.08)
Texas 1.71 70.72 0.111 0.93 (0.03) −0.09 (0.04) −0.05 (0.01) 1.17 (0.20) −0.98 (0.05) 0.65 (0.07) −0.64 (0.07)

A .12.4 Count of Walks Centrality


A .12 L E A R N E D PA R A M E T E R S O F D I F F E R E N T C E N T R A L I T Y B A S E D G S O S 137

Table a.11 – Graph Properties of the used datasets and the corresponding learned hyperparameters in GAGCN w/
K-Core
Dataset Graph Properties Hyperparameters

Avg. K-core Avg. Count. Walks homophily e1 e2 e3 m1 m2 m3 a

Physics 7.71 449.22 0.931 0.38 (0.03) −0.35 (0.01) −0.35 (0.01) 0.34 (0.02) 1.28 (0.01) 0.30 (0.02) 1.35 (0.02)
Photo 16.97 3204.098 0.827 0.52 (0.02) −0.31 (0.01) −0.31 (0.01) 0.73 (0.03) 1.44 (0.02) 0.59 (0.03) 1.77 (0.05)
Cora 2.31 42.52 0.809 0.34 (0.01) −0.74 (0.00) 0.24 (0.01) 0.60 (0.01) 1.55 (0.01) 0.59 (0.01) 0.68 (0.01)
CS 4.94 162.75 0.808 0.41 (0.01) −0.29 (0.01) −0.29 (0.01) 0.43 (0.01) 1.39 (0.01) 0.39 (0.01) 1.47 (0.01)
PubMed 2.39 75.43 0.802 0.27 (0.00) −0.31 (0.00) −0.32 (0.00) 0.34 (0.01) 1.34 (0.01) 0.34 (0.01) 1.35 (0.01)
Computers 18.84 6221.39 0.777 0.51 (0.02) −0.28 (0.01) −0.29 (0.01) 0.78 (0.03) 1.50 (0.01) 0.66 (0.03) 1.72 (0.05)
CiteSeer 1.73 18.91 0.735 0.39 (0.01) −0.26 (0.00) −0.26 (0.00) 0.45 (0.01) 1.43 (0.01) 0.44 (0.01) 1.46 (0.00)
ogbn-arxiv 7.13 4898.16 0.654 0.27 (0.01) −0.53 (0.01) −0.55 (0.01) −0.70 (0.01) −1.04 (0.03) 0.31 (0.02) 0.70 (0.02)
deezer-europe 3.57 106.16 0.525 −0.01 (0.03) −0.51 (0.00) −0.51 (0.00) 0.02 (0.01) 0.99 (0.01) 0.02 (0.01) 1.07 (0.01)
Penn94 33.68 10662.08 0.470 −0.09 (0.30) −0.39 (0.08) −0.40 (0.08) 0.28 (0.36) 1.27 (0.16) 0.05 (0.41) 1.60 (0.15)
chameleon 16.60 2913.48 0.231 0.15 (0.04) −0.06 (0.01) −0.06 (0.01) −0.17 (0.02) 0.88 (0.01) −0.16 (0.02) −0.15 (0.01)
squirrel 41.55 31888.02 0.222 0.46 (0.02) −0.78 (0.01) 0.24 (0.01) −0.97 (0.05) 1.78 (0.04) −0.97 (0.05) −1.08 (0.12)
arxiv-year 7.13 82.85 0.218 0.34 (0.05) −0.36 (0.01) −0.41 (0.01) −0.13 (0.06) 1.03 (0.01) −0.03 (0.02) 0.80 (0.02)
Wisconsin 2.05 76.26 0.206 1.20 (0.02) −0.05 (0.02) −0.06 (0.02) 1.48 (0.03) −0.96 (0.04) 0.54 (0.02) −0.54 (0.03)
Cornell 1.74 58.47 0.132 0.34 (0.05) −0.36 (0.01) −0.41 (0.01) −0.13 (0.06) 1.03 (0.01) −0.03 (0.02) 0.80 (0.02)
Texas 1.71 70.72 0.111 0.50 (0.06) −0.02 (0.02) −0.04 (0.02) −0.87 (0.05) 1.04 (0.05) −0.85 (0.05) −0.86 (0.05)

Table a.12 – Graph Properties of the used datasets and the corresponding learned hyperparameters in GAGCN w/
PageRank
Dataset Graph Properties Hyperparameters

Avg. K-core Avg. Count. Walks homophily e1 e2 e3 m1 m2 m3 a

Physics 7.71 449.22 0.931 0.51 (0.00) 0.00 (0.00) 0.00 (0.00) 1.31 (0.06) 1.00 (0.08) 0.34 (0.06) 0.33 (0.07)
Photo 16.97 3204.098 0.827 0.53 (0.02) 0.08 (0.01) 0.08 (0.01) 0.88 (0.01) 0.85 (0.01) −0.12 (0.01) −0.13 (0.01)
Cora 2.31 42.52 0.809 0.00 (0.00) −0.71 (0.02) 0.11 (0.01) 0.63 (0.01) 1.49 (0.02) 0.63 (0.01) 0.67 (0.01)
CS 4.94 162.75 0.808 0.00 (0.00) −0.10 (0.01) −0.10 (0.01) 0.46 (0.04) 1.38 (0.10) 0.46 (0.04) 1.49 (0.03)
PubMed 2.39 75.43 0.802 0.51 (0.00) 0.00 (0.00) 0.00 (0.00) 1.34 (0.02) 1.28 (0.02) 0.36 (0.02) 0.38 (0.02)
Computers 18.84 6221.39 0.777 0.00 (0.00) −0.42 (0.00) −0.42 (0.00) −0.13 (0.02) 0.84 (0.01) −0.13 (0.02) 0.87 (0.02)
CiteSeer 1.73 18.91 0.735 0.00 (0.00) −0.14 (0.00) −0.12 (0.00) 0.44 (0.01) 1.41 (0.00) 0.44 (0.01) 1.47 (0.01)
ogbn-arxiv 7.13 4898.16 0.654 0.00 (0.00) −0.89 (0.01) 0.11 (0.01) −0.22 (0.01) 0.78 (0.01) −0.22 (0.01) −0.22 (0.01)
deezer-europe 3.57 106.16 0.525 0.57 (0.05) 0.05 (0.01) 0.05 (0.01) 0.90 (0.03) 0.90 (0.02) −0.10 (0.03) −0.10 (0.02)
Penn94 33.68 10662.08 0.470 0.54 (0.01) 0.05 (0.01) 0.05 (0.01) 1.10 (0.01) −0.90 (0.01) 0.10 (0.01) −0.10 (0.01)
chameleon 16.60 2913.48 0.231 −0.01 (0.00) −0.94 (0.00) 0.06 (0.01) 0.27 (0.05) -0.88 (0.02) 1.26 (0.05) −0.25 (0.05)
squirrel 41.55 31888.02 0.222 0.00 (0.00) −0.43 (0.01) −0.43 (0.01) 0.14 (0.01) −0.86 (0.01) 1.14 (0.01) −0.14 (0.01)
arxiv-year 7.13 82.85 0.218 0.00 (0.00) −0.84 (0.03) 0.06 (0.01) −0.04 (0.02) 0.86 (0.03) −0.04 (0.02) −0.05 (0.03)
Wisconsin 2.05 76.26 0.206 0.64 (0.02) 0.10 (0.01) 0.10 (0.01) 1.68 (0.03) −1.01 (0.04) 0.69 (0.02) −0.69 (0.02)
Cornell 1.74 58.47 0.132 0.64 (0.03) 0.11 (0.01) 0.11 (0.01) 1.71 (0.02) −1.03 (0.05) 0.71 (0.01) −0.72 (0.01)
Texas 1.71 70.72 0.111 0.68 (0.03) 0.14 (0.03) 0.14 (0.03) 1.63 (0.04) −0.93 (0.06) 0.64 (0.04) −0.63 (0.04)
138 A P P E N D I X : R E T H I N K I N G G S O S I N G N N S F O R G R A P H R E P R E S E N TAT I O N L E A R N I N G

Table a.13 – Graph Properties of the used datasets and the corresponding learned hyperparameters in GAGCN w/
Count of walks.
Dataset Graph Properties Hyperparameters

Avg. K-core Avg. Count. Walks homophily e1 e2 e3 m1 m2 m3 a

Physics 7.71 449.22 0.931 0.95 (0.01) −0.02 (0.02) −0.02 (0.02) 0.89 (0.02) 0.96 (0.03) −0.10 (0.01) −0.09 (0.01)
Photo 16.97 3204.098 0.827 −0.06 (0.01) −0.07 (0.00) −0.07 (0.00) −0.05 (0.01) 0.87 (0.00) −0.04 (0.02) −0.09 (0.01)
Cora 2.31 42.52 0.809 0.36 (0.02) 0.03 (0.01) 0.02 (0.01) 0.68 (0.02) 1.37 (0.09) 0.63 (0.02) 0.64 (0.01)
CS 4.94 162.75 0.808 0.45 (0.02) 0.03 (0.01) 0.02 (0.01) 0.62 (0.03) 1.23 (0.04) 0.47 (0.02) 0.52 (0.02)
PubMed 2.39 75.43 0.802 0.30 (0.02) −0.15 (0.02) −0.16 (0.02) 0.60 (0.02) 1.58 (0.03) 0.56 (0.02) 1.38 (0.04)
Computers 18.84 6221.39 0.777 −0.05 (0.01) −0.07 (0.00) −0.07 (0.00) −0.05 (0.02) 0.86 (0.01) −0.04 (0.03) −0.10 (0.02)
CiteSeer 1.73 18.91 0.735 0.25 (0.01) −0.13 (0.01) −0.14 (0.01) 0.67 (0.01) 1.63 (0.01) 0.62 (0.01) 1.63 (0.01)
ogbn-arxiv 7.13 4898.16 0.654 0.12 (0.00) −0.08 (0.00) −0.19 (0.00) 0.32 (0.00) 1.57 (0.02) 0.32 (0.00) 0.93 (0.01)
deezer-europe 3.57 106.16 0.525 0.26 (0.03) −1.04 (0.03) −0.07 (0.03) 0.58 (0.04) 0.88 (0.06) 0.60 (0.04) 0.67 (0.06)
Penn94 33.68 10662.08 0.470 0.30 (0.00) −0.77 (0.03) 0.06 (0.09) 0.58 (0.00) −0.46 (0.16) 1.39 (0.01) −0.31 (0.17)
chameleon 16.60 2913.48 0.231 0.32 (0.06) −0.05 (0.01) −0.05 (0.01) −0.28 (0.08) 0.89 (0.02) −0.17 (0.03) −0.14 (0.02)
squirrel 41.55 31888.02 0.222 0.23 (0.11) −0.71 (0.10) 0.35 (0.10) −0.46 (0.46) 1.28 (0.22) −0.32 (0.33) −0.29 (0.48)
arxiv-year 7.13 82.85 0.218 0.23 (0.01) −0.28 (0.01) −0.16 (0.01) −0.26 (0.01) 0.99 (0.03) −0.01 (0.04) 0.84 (0.03)
Wisconsin 2.05 76.26 0.206 0.40 (0.01) −0.79 (0.07) 0.18 (0.05) 0.61 (0.01) −1.38 (0.08) 1.48 (0.01) −0.69 (0.06)
Cornell 1.74 58.47 0.132 0.40 (0.01) −0.21 (0.04) −0.22 (0.04) 0.59 (0.01) −1.51 (0.06) 1.47 (0.01) −0.61 (0.05)
Texas 1.71 70.72 0.111 0.40 (0.01) −0.72 (0.03) 0.24 (0.03) 0.58 (0.01) −1.48 (0.05) 1.47 (0.01) −0.59 (0.03)
A P P E N D I X : A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N
b

139
140 A P P E N D I X : A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

B .1 D ATA S E T S TAT I S T I C S

Characteristics and information about the datasets utilized in the node classification part of the study
are presented in Table b.1, which provides statistics about each dataset, and highlights the variety in their
features, size, and edge homophily,

Table b.1 – Statistics of the node classification datasets used in our experiments.

Dataset #Features #Nodes #Edges #Classes Edge Homophily

Cora 1,433 2,708 5,208 7 0.809


CiteSeer 3,703 3,327 4,552 6 0.735
PubMed 500 19,717 44,338 3 0.802
CS 6,805 18,333 81,894 15 0.808
chameleon 2,325 2,277 62,792 5 0.231
Cornell 1,703 183 557 5 0.132
squirrel 2,089 5,201 396,846 5 0.222
Wisconsin 1,703 251 916 5 0.206
Texas 1,703 183 574 5 0.111
Photo 745 7,650 238,162 8 0.827
ogbn-arxiv 128 169,343 2,315,598 40 0.654
Computers 767 13752 491,722 10 0.777

B .2 N O D E - S P E C I F I C D E P T H A N A LY S I S I N G R A P H N E U R A L N E T W O R K S

We generate synthetic graphs of size N = 5, 000. We select nodes belonging to sparse or dense
regions in the original graph based on their core number. We consider only nodes with labels that are
sufficiently present in both dense and sparse region. Last, we randomly select nodes, all by keeping the
label distribution similar in both sparse and dense subgraphs. Yellow points indicate edges, while purple
points represent non-edges. Notably, there are variations in edge density across different blocks: the first
block is characterized by extreme sparsity, whereas the second block exhibits a much denser structure.

B .3 TIME COMPLEXITY

Understanding the time complexity of the ADMP-GCN model is crucial for evaluating its practical efficiency
and scalability. Table b.2 reports the average training time, measured in seconds, for two distinct settings:
ALM and ST.
B .4 H Y P E R PA R A M E T E R C O N F I G U R AT I O N S 141

0 Computers 0 Photo
200 100
200
400
300
600 400
800 500
600
0 250 500 750 0 200 400 600
Figure b.1 – The adjacency matrix of the synthetic graphs extracted from the real graphs Computers and Photo.

Table b.2 – The average time needed for each training setting for different datasets.

Model Cora squirel chamelon Computers Photo Ogbn-arxiv

ADMP-GCN ALM 14 36 15 51 26 266


ADMP-GCN ST 32 88 40 175 87 882

B .4 H Y P E R PA R A M E T E R C O N F I G U R AT I O N S

For a more balanced comparison, however, we use the same training procedure for all the models. The
hyperparameters in each dataset where performed using a Grid search on the classical GCN over the
following search space:
— Hidden size: [16, 32, 64, 128, 256, 512],,
— Learning rate: [0.1, 0.01, 0.001],
— Dropout probability: [16, 32, 64, 128, 256, 512].,
The number of layers was fixed to 2. The optimal hyperparameters can be found in Table b.3.

B .5 S U P P L E M E N TA R Y R E S U LT S O F A D M P - G C N

We provide additional results on the ADMP-GCN in Tables b.4, b.5 and b.6.

B .6 S U P P L E M E N TA R Y R E S U LT S O F A D M P - G I N

Further results on the ADMP-GIN can be found in Tables b.7, b.8, b.9, b.10 and b.11.
142 A P P E N D I X : A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

Table b.3 – Hyperparameters used in our experiments.

Dataset Hidden Size Learning Rate Dropout Probability

Cora 64 0.01 0.8


CiteSeer 64 0.01 0.4
PubMed 64 0.01 0.2
CS 512 0.01 0.4
arxiv-year 512 0.01 0.2
chameleon 512 0.01 0.2
cornell 512 0.01 0.2
deezer-europe 512 0.01 0.2
squirrel 512 0.01 0.2
Wisconsin 512 0.01 0.2
Texas 512 0.01 0.2
Photo 512 0.01 0.6
OGBN-Arxiv 512 0.01 0.5
Computers 512 0.01 0.2
Physics 512 0.01 0.4
Penn94 64 0.01 0.2
B .6 S U P P L E M E N TA R Y R E S U LT S O F A D M P - G I N 143

Table b.4 – Comparison of ADMP-GCN training paradigms ALM and ST. These paradigms are also compared to
the single-task training setting to evaluate which approach most closely mimics the classical GCN under
single-task training. The best results for each dataset are bolded.
#Layers Training Paradigm Model Photo Computers Chameleon Cornell Wisconsin Texas Squirrel
Single-task GCN 70.18±3.39 56.83±2.92 33.55±0.00 40.54±0.00 70.59±0.00 64.86±0.00 26.42±0.00
0
ADMP-GCN (ALM) 68.23±3.11 56.61±2.07 30.70±0.00 40.54±0.00 52.94±0.00 64.86±0.00 22.60±0.06
Multi-task
ADMP-GCN (ST) 70.56±3.33 57.94±3.62 33.55±0.00 40.54±0.00 70.59±0.00 64.86±0.00 26.42±0.00
Single-task GCN 65.86±5.18 59.62±3.40 38.16±0.00 40.54±0.00 56.86±0.00 64.86±0.00 21.23±0.00
1
ADMP-GCN (ALM) 64.49±6.95 57.60±6.63 27.41±0.00 40.54±0.00 54.90±0.00 64.86±0.00 19.31±0.00
Multi-task
ADMP-GCN (ST) 64.03±3.63 56.61±5.44 38.16±0.00 40.54±0.00 56.86±0.00 64.86±0.00 21.23±0.00
Single-task GCN 82.32±2.97 69.05±2.60 58.73±0.96 54.59±2.91 57.06±1.85 62.16±1.21 33.40±0.68
2
ADMP-GCN (ALM) 48.10±6.28 34.06±10.87 40.39±0.95 44.32±1.32 56.86±1.52 62.16±0.00 21.73±0.28
Multi-task
ADMP-GCN (ST) 85.80±0.43 68.22±4.53 58.77±1.08 55.95±4.02 56.86±1.52 61.89±1.46 33.57±0.46
Single-task GCN 86.51±2.33 75.32±4.18 58.40±2.46 45.95±3.42 43.33±2.23 44.59±4.05 36.29±0.73
3
ADMP-GCN (ALM) 79.78±4.17 53.40±9.51 49.67±1.03 38.92±5.43 46.47±2.16 50.00±3.25 29.65±2.22
Multi-task
ADMP-GCN (ST) 85.82±1.76 73.69±3.67 50.79±1.02 48.38±3.07 47.25±2.23 57.30±5.38 32.95±0.45
Single-task GCN 75.36±12.68 61.77±20.27 51.12±7.56 47.03±2.16 51.57±2.91 48.38±3.51 37.73±0.80
4
ADMP-GCN (ALM) 82.37±4.86 62.61±5.71 49.63±1.90 44.59±4.23 48.04±2.67 52.43±4.86 28.99±1.40
Multi-task
ADMP-GCN (ST) 87.08±0.77 72.35±5.44 53.86±2.44 51.35±3.20 50.39±1.76 57.84±5.01 34.74±0.48
Single-task GCN 78.49±7.79 55.63±14.13 50.07±1.73 46.76±4.53 45.88±2.18 53.51±6.02 36.52±1.01
5
ADMP-GCN (ALM) 82.57±5.33 63.05±9.21 48.36±1.86 48.92±3.72 46.67±4.71 49.46±5.93 29.60±1.70
Multi-task
ADMP-GCN (ST) 86.37±0.85 74.53±3.65 53.86±1.39 52.70±3.02 50.59±1.47 56.49±3.07 32.11±1.14

Table b.5 – Comparison of highest accuracy (± standard deviation) for GCN and ADMP-GCN ST, with the layer
achieving the best accuracy indicated in brackets. The final row shows the Oracle Accuracy for ADMP-
GCN ST. Best results for each dataset are bolded.
Model Photo Computers chamelon Cornell Wisconsin Texas Squirrel

GCN 86.51±2.33 [3] 75.32±4.18 [3] 58.73±0.96 [2] 54.59±2.91 [2] 70.59±0.00 [0] 64.86±0.00 [0] 37.73±0.80 [4]
ADMP-GCN DT 87.08±0.77 [4] 74.53±3.65 [5] 58.77±1.08 [2] 55.95±4.02 [2] 70.59±0.00 [0] 64.86±0.00 [0] 34.74±0.48 [4]
ADMP-GCN ST - Oracle 96.14±0.81 90.15±0.97 79.50±1.31 74.59±3.24 80.39±0.00 77.84±2.36 69.54±0.65
144 A P P E N D I X : A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

Table b.6 – Comparison of ADMP-GCN training paradigms ALM and ST. These paradigms are also compared to
the single-task training setting to evaluate which approach most closely mimics the classical GCN under
single-task training. The best results for each dataset are bolded.
Model Photo Computers chamelon Cornell Wisconsin Texas Squirrel

JKNET-CAT 87.92±1.98 [2] 74.68±6.92 [3] 56.89±1.77 [2] 46.22±5.73 [5] 72.55±0.00 [0] 64.86±0.00 [0] 41.26±0.88 [2]
JKNET-MAX 88.02±2.21 [2] 77.97±2.57 [0] 57.70±2.79 [2] 45.95±5.41 [5] 62.75±2.63 [1] 68.65±4.05 [5] 41.36±0.59 [2]
JKNET-LSTM 87.74±1.94 [0] 77.13±2.80 [1] 53.07±4.47 [2] 43.78±3.78 [2] 62.94±2.23 [1] 64.86±0.00 [0] 41.32±0.58 [5]
Residuals - GCNII 86.99±2.47 [2] 62.83±19.61 [2] 61.03±2.07 [2] 55.14±8.48 [5] 70.59±0.00 [0] 64.86±0.00 [0] 35.90±1.25 [5]
AdaGCN 64.84±0.89 58.68±1.31 47.37±0.48 70.00±4.75 75.88±1.76 72.43±1.62 29.15±0.77
GPR-GNN 89.16±2.04 [2] 77.43±3.32 [2] 63.29±0.91 [2] 62.70±4.65 [2] 58.82±1.24 [2] 59.19±3.72 [2] 38.52±1.06 [4]

GCN 86.51±2.33 75.32±4.18 58.73±0.96 54.59±2.91 70.59±0.00 64.86±0.00 37.73±0.80


ADMP-GCN 87.08±0.77 74.53±3.65 58.77±1.08 55.95±4.02 70.59±0.00 64.86±0.00 34.74±0.48

ADMP-GCN w/ Degree 88.15±1.52 75.53±2.09 58.75±0.91 44.32±4.22 63.92±1.30 56.49±1.89 35.07±1.13


ADMP-GCN w/ k-core 88.31 ±1.31 75.57±1.88 58.97±1.09 48.11±4.65 68.82±1.37 60.81±3.25 34.46±0.70
ADMP-GCN w/ Walk Count 88.42±1.51 76.14 ±2.08 58.29±1.22 51.89±4.32 67.45±1.80 68.65±2.16 34.39±0.89
ADMP-GCN w/ PageRank 88.21±1.45 75.50±2.23 58.44±1.23 55.14 ±4.86 65.29±2.16 60.00±3.15 34.68 ±0.66

Table b.7 – Comparison of ADMP-GIN training paradigms ALM and ST. These paradigms are also compared to
the single-task training setting to evaluate which approach most closely mimics the classical GIN under
single-task training. The best results for each dataset are bolded.
# Layers Training Paradigm Model Cora CiteSeer CS PubMed Genius Ogbn-arxiv
Single-task GIN 56.40±0.06 57.14±0.09 87.17±1.41 72.49±0.08 80.78±1.03 48.87±0.04
0
ADMP-GIN (ALM) 58.00±0.00 61.50±0.00 86.36±0.78 73.20±0.00 79.95±0.10 36.49±0.19
Multi-task
ADMP-GIN (ST) 56.38±0.04 57.17±0.08 87.41±0.95 72.47±0.14 80.47±0.91 48.87±0.04
Single-task GIN 75.17±0.09 64.79±0.03 90.29±0.99 74.97±0.11 78.42±4.95 60.90±0.15
1
ADMP-GIN (ALM) 74.50±0.00 66.50±0.00 88.66±1.18 75.40±0.00 72.07±17.44 59.43±0.73
Multi-task
ADMP-GIN (ST) 75.07±0.06 64.80±0.00 90.82±1.15 75.00±0.00 77.01±12.06 60.85±0.01
Single-task GIN 77.73±0.99 65.23±1.45 87.93±0.71 76.05±1.14 78.89±0.48 16.23±9.60
2
ADMP-GIN (ALM) 63.53±4.80 60.68±2.09 12.90±8.05 72.20±4.11 79.60±0.51 47.63±3.20
Multi-task
ADMP-GIN (ST) 78.07±0.68 65.41±1.91 87.47±2.18 76.04±0.88 72.99±13.48 10.13±8.00
Single-task GIN 74.40±1.14 60.81±2.20 82.13±2.24 74.97±1.69 52.22±28.47 6.00±0.17
3
ADMP-GIN (ALM) 66.99±2.85 59.57±2.57 17.69±12.02 74.15±2.00 68.00±23.98 27.51±16.25
Multi-task
ADMP-GIN (ST) 76.28±1.11 65.13±1.00 83.60±3.72 76.21±1.75 80.04±0.09 13.74±9.66
Single-task GIN 67.68±4.07 57.54±2.92 47.37±17.10 73.62±1.63 80.05±0.10 6.07±0.21
4
ADMP-GIN (ALM) 68.78±4.30 60.12±2.36 21.07±14.15 74.20±2.57 79.36±1.18 15.23±11.85
Multi-task
ADMP-GIN (ST) 74.94±1.58 65.21±1.54 81.33±2.15 76.46±1.04 80.04±0.09 16.02±10.78
Single-task GIN 32.48±10.47 54.76±2.32 18.56±11.32 67.98±5.80 80.05±0.10 6.19±0.39
5
ADMP-GIN (ALM) 67.46±4.34 59.74±1.95 29.16±13.84 73.94±2.42 79.99±0.10 14.61±12.11
Multi-task
ADMP-GIN (ST) 71.34±2.09 63.89±1.39 79.10±1.84 75.75±1.09 79.57±1.40 15.72±12.13
B .6 S U P P L E M E N TA R Y R E S U LT S O F A D M P - G I N 145

Table b.8 – Comparison of ADMP-GIN training paradigms ALM and ST. These paradigms are also compared to
the single-task training setting to evaluate which approach most closely mimics the classical GIN under
single-task training. The best results for each dataset are bolded.
#Layers Training Paradigm Model Photo Computers Chameleon Cornell Wisconsin Texas Squirrel
Single-task GIN 70.19±2.91 56.83±3.89 33.55±0.00 40.54±0.00 70.59±0.00 64.86±0.00 26.42±0.00
0
ADMP-GIN ALM 69.16±3.27 56.27±3.00 30.70±0.00 40.54±0.00 52.94±0.00 64.86±0.00 22.62±0.05
Multi-task
ADMP-GCN ST 68.68±3.36 56.71±2.77 33.55±0.00 40.54±0.00 70.59±0.00 64.86±0.00 26.42±0.00
Single-task GIN 82.32±2.03 70.90±2.64 61.03±0.10 40.54±0.00 56.86±0.00 64.86±0.00 47.65±0.27
1
ADMP-GIN ALM 78.79±2.84 65.28±2.17 56.75±0.09 40.54±0.00 54.90±0.00 64.86±0.00 44.12±0.39
Multi-task
ADMP-GCN ST 83.35±1.67 71.88±3.80 60.96±0.00 40.54±0.00 56.86±0.00 64.86±0.00 47.65±0.00
Single-task GIN 83.86±2.19 63.39±10.30 63.57±1.01 61.62±2.65 54.71±2.70 64.86±2.96 19.84±1.59
2
ADMP-GIN ALM 25.83±9.50 13.04±8.60 22.37±0.00 40.00±3.97 49.61±3.51 64.32±2.91 20.11±0.85
Multi-task
ADMP-GCN ST 68.44±23.78 42.65±17.47 62.30±2.18 59.46±2.42 53.33±2.60 64.32±1.62 19.31±0.00
Single-task GIN 28.63±13.64 19.90±9.20 26.29±1.94 35.68±2.65 48.04±4.04 60.54±5.82 24.84±2.06
3
ADMP-GIN ALM 17.94±6.70 23.04±13.93 26.29±1.31 44.05±4.53 50.20±5.13 64.59±7.69 19.95±2.40
Multi-task
ADMP-GCN ST 71.32±19.80 54.40±16.24 48.84±2.76 43.51±3.51 56.27±5.26 64.59±14.58 26.03±0.52
Single-task GIN 14.43±4.93 12.51±8.60 29.63±2.48 43.78±4.15 51.57±1.53 66.76±4.69 20.47±1.26
4
ADMP-GIN ALM 11.81±6.16 9.01±9.74 21.01±2.16 45.14±4.02 50.78±3.22 64.05±3.83 19.18±1.57
Multi-task
ADMP-GCN ST 68.49±16.64 52.61±19.82 44.54±4.16 45.14±4.84 50.59±3.14 68.92±5.70 26.59±1.36
Single-task GIN 14.22±5.42 12.55±7.97 26.47±3.22 43.51±4.43 50.20±6.09 66.22±1.81 20.60±1.25
5
ADMP-GIN ALM 15.84±4.93 9.81±7.59 28.82±5.47 42.97±5.33 49.41±4.09 65.14±5.05 19.77±1.52
Multi-task
ADMP-GCN ST 68.48±13.80 58.96±9.83 43.31±3.01 44.05±4.53 45.69±5.19 67.57±9.44 29.74±2.46

Table b.9 – Comparison of highest accuracy (± standard deviation) for GIN and ADMP-GIN ST, with the layer
achieving the best accuracy indicated in brackets. The final row shows the Oracle Accuracy for ADMP-
GIN ST. Best results for each dataset are bolded.
Model Cora CiteSeer CS PubMed Genius Ogbn-arxiv

GIN 77.73±0.99 [2] 65.23±1.45 [2] 90.29±0.99 [1] 76.05±1.14 [2] 80.78±1.03 [0] 60.90±0.15 [1]
ADMP-GIN (ST) 78.07±0.68 [2] 65.41±1.91 [2] 90.82±1.15 [1] 76.46±1.04 [4] 80.47±0.91 [0] 60.85±0.01 [1]
ADMP-GIN ST - Oracle 90.76±0.27 81.87±0.63 97.64±0.18 92.73±0.54 92.07±6.13 71.23±2.69

Table b.10 – Comparison of highest accuracy (± standard deviation) for GIN and ADMP-GIN ST, with the layer
achieving the best accuracy indicated in brackets. The final row shows the Oracle Accuracy for ADMP-
GIN ST. Best results for each dataset are bolded.
Model Photo Computers chamelon Cornell Wisconsin Texas Squirrel

GIN 83.86±2.19 [2] 70.90±2.64 [1] 63.57±1.01 [2] 61.62±2.65 [2] 70.59±0.00 [0] 66.76±4.69 [4] 47.65±0.27 [1]
ADMP-GIN ST 83.35±1.67 [1] 71.88±3.80 [1] 62.30±2.18 [2] 59.46±2.42 [2] 70.59±0.00 [0] 68.92±5.70 [4] 47.65±0.00 [1]
ADMP-GIN ST - Oracle 95.92±1.03 90.40±2.66 86.73±0.90 73.24±4.75 80.78±1.18 84.05±3.07 78.41±0.89
146 A P P E N D I X : A D A P T I V E D E P T H M E S S A G E PA S S I N G G N N

Table b.11 – Classification accuracy (± standard deviation) on different node classification datasets for the baselines
based on the GIN backbone. The higher the accuracy (in %), the better the model. Highlighted are the
first, second best results. OOM means Out of memory.
Model Photo Computers chamelon Cornell Wisconsin Texas squirrel

JKNET-MAX 82.48±2.95 [1] 69.62±3.57 [1] 59.61±2.24 [2] 47.84±4.84[4] 57.84±2.19 [0] 66.49±2.16 [5] 23.19±3.55 [2]
JKNET-LSTM 87.02±1.42 [2] 76.50±1.44 [2] 62.61±1.39 [2] 50.54±3.64 [2] 57.84±3.64[4] 70.81±3.15 [1] 26.34±4.31 [0]
GPR-GIN 85.82±1.02 [2] 76.01±3.20 [3] 67.26±0.50 [2] 62.97±2.97 [5] 68.04±3.62[4] 73.78±2.72 [5] 46.78±1.80 [2]

GIN 83.86±2.19 [2] 70.9±2.64 [1] 63.57±1.01 [2] 61.62±2.65 [2] 70.59±0.00 [0] 66.76±4.69[4] 47.65±0.27 [1]
ADMP-GIN 83.35±1.67 [1] 71.88±3.8 [1] 62.30±2.18 [2] 59.46±2.42 [2] 70.59±0.00 [0] 68.92±5.70[4] 47.65±0.00 [1]

ADMP-GIN w/ Degree 84.15±1.17 74.55±2.89 64.43±1.47 46.49±4.65 66.47±2.83 69.46±3.21 47.65±0.00


ADMP-GIN w/ k-core 84.08±1.09 74.58±3.56 64.65±1.45 46.76±4.84 70.78±4.68 71.89±4.22 47.65±0.00
ADMP-GIN w/ Walk Count 84.03±1.31 73.85±3.70 63.38±1.53 58.11±6.07 63.14±3.90 74.86±2.97 47.65±0.00
ADMP-GIN w/ PageRank 84.36±1.32 74.27±2.97 63.31±1.10 55.41±3.68 65.49±3.3 70.81±3.78 47.65±0.00
A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S
c

147
148 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

C.1 PROOF OF LEMMA 5.3.2

α,β
Proof. Let Wϵ = {( Aw , Xw ); dα,β ([ G, X ], [ G̃, X̃ ]) ≤ ϵ} be the set of “worst-case” adversarial attacks within
the attack budget ϵ. We denote the expected vulnerability of a graph function f on the set of worst-case
α,β
adversarial examples as Advϵ [ f ]|W α,β .
ϵ
α,β
By definition, we have Wϵ ⊂ Bα,β ( G, X, ϵ). Consequently, we have that

Advϵ [ f ]|W α,β = P(G,X )∼DG ,X [( G̃, X̃ ) ∈ Wϵ


α,β α,β
: dY ( f ( G̃, X̃ ), f ( G, X )) > σ ] (c.1)
ϵ

≤ P(G,X )∼DG ,X [( G̃, X̃ ) ∈ Bα,β ( G, X, ϵ) : dY ( f ( G̃, X̃ ), f ( G, X )) > σ] (c.2)


α,β
≤ Advϵ [ f ]. (c.3)

α,β
Let the graph function f be ((dα,β , ϵ), (dY , γ))–robust, then Advϵ [ f ] ≤ γ. From c.3, we have
α,β α,β
W − Advϵ [ f ] ≤ Advϵ [ f ] ≤ γ.
We therefore, can conclude that if f is ((dα,β , ϵ), (dY , γ))–robust, then f is ((dα,β , ϵ), (dY , γ)) − “worst-case” robust.

C.2 PROOF OF PROPOSITION C.2.1

Proposition c.2.1. Let f : (G , X ) → Y be a graph-based classifier subject to feature-based attacks and X


be of dimension D. Let d0,1 0,1
2 be the graph distance corresponding to the spectral norm, d1 corresponding
0,1 0,1
to the L1 norm and d p corresponding to the L p norm with p > 2. If f is ((d2 , ϵ), (dY , γ))–robust, then f is

also ((d0,1
p ,D
−1/2 ϵ ) , ( d , γ ))–robust and (( d0,1 , D −2 ϵ ) , ( d , Dγ ))–robust.
Y 1 Y

Proof. Let RD the real finite-dimensional space. The set of standard considered norms of RD for x =
( x1 , . . . , x D ) ∈ RD is
!1/p
D
∀ p > 0, ∥ x ∥ p = ∑ | xi | p , ∥ x ∥∞ = max | xi |.
1≤ i ≤ n
i =1

As for any x ∈ RD , the mapping p 7→ ∥ x ∥ p is monotone decreasing (Raıssouli and Jebril, 2010) , meaning
that
q > p ≥ 1 ⇒ ∥ x ∥q ≤ ∥ x ∥ p .
Using Holder’s inequality with s = q/p > 1, we have
D
∑ | xi | p
p
∥x∥ p =
i =1
D
= ∑ | xi | p · 1
i =1
! 1s !1− 1s
D D
∑ (| xi | ∑1
s
p s
≤ ) s −1

i =1 i =1
! qp !1− qp
D D q

∑ (| xi | p ) ∑1
q/p
= q− p

i =1 i =1
1− p/q
p
=D ∥ x ∥q .
C.2 P R O O F O F P R O P O S I T I O N C . 2 . 1 149

Thus,
q > p ≥ 1 ⇒ ∥ x ∥ p ≤ D1/p−1/q ∥ x ∥q . (c.4)
We additionally have that
!1/q
D
∥ x ∥q = ∑ | xi | q
,
i =1
!1/q
D

q
≤ ∥ x ∥∞
i =1
1/q
≤D ∥ x ∥∞
≤ D1/q ∥ x ∥ p . (c.5)

From c.4 and c.5, we deduce that


1
− 1p 1
q > p ≥ 1 ⇒ Dq ∥ x ∥ p ≤ ∥ x ∥q ≤ D q ∥ x ∥ p ,

Consequently, for any matrix N ∈ Rn× D , we have the following


( 1 1 1

D q p ∥ Nx ∥ p ≤ ∥ Nx ∥q ≤ D q ∥ Nx ∥ p ,
p > q ≥ 1 ⇒ ∀ x ̸= 0 1
−1 1
D q p ∥ x ∥ p ≤ ∥ x ∥q ≤ D q ∥ x ∥ p ,
 1 1
 D q − p ∥ Nx ∥ ≤ ∥ Nx ∥ ≤ D 1q ∥ Nx ∥ ,
p q p
⇒ ∀ x ̸= 0 − 1
1 1 − 1
+ 1
1
 D q ≤ ≤D q p
∥x∥ p
,
∥ x ∥q ∥x∥ p
1 ∥ Nx ∥ p ∥ Nx ∥q 1 ∥ Nx ∥ p
⇒ ∀ x ̸= 0, D − p ≤ ≤ Dp ,
∥x∥ p ∥ x ∥q ∥x∥ p
1 1
⇒ D − p ∥ N ∥ p ≤ ∥ N ∥q ≤ D p ∥ N ∥ p .

In particular, for p = 2, we have



∀q > 2, ∀ N ∈ Rn× D ∥ N ∥2 ≤ D ∥ N ∥q . (c.6)

Similar for q = 2 and p = 1, we have

∀ N ∈ Rn × D ∥ N ∥2 ≤ D ∥ N ∥1 . (c.7)

The Inequalities c.6 and c.7 stay valid for the distance related to the norm L p , i. e.

∀ X, X̃ ∈ Rn× D , ∀ p > 2, d2 ( X, X̃ ) ≤ Dd p ( X, X̃ ),
d2 ( X, X̃ ) ≤ Dd1 ( X, X̃ ).

Finally, for every input graph G, X, we have


1
∀ p > 2, {( G ′ , X ′ ) | d p ([ G, X ], [ G̃, X̃ ]) < √ ϵ} ⊂ {( G ′ , X ′ ) | d2 ([ G, X ], [ G̃, X̃ ]) < ϵ},
D
1
{( G ′ , X ′ ) | d1 ([ G, X ], [ G̃, X̃ ]) < ϵ} ⊂ {( G ′ , X ′ ) | d1 ([ G, X ], [ G̃, X̃ ]) < ϵ}.
D
Therefore, the ((d0,1 -robustness of a graph function f implies the ((d0,1
2 , ϵ ) , ( dY , γ ))√ p ,D
−1/2 ϵ ) , ( d , γ ))–
Y
robustness and ((d0,1
1 , D −1 ϵ ) , ( d , Dγ ))–robustness.
Y
150 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

C.3 PROOF OF THEOREM 5.4.1 AND 5.4.2

Proof. Let’s consider a graph-function f that is based on L layers of GCN. The GCN message-passing
propagation can be written for a node u as
(ℓ−1)
(ℓ) W (ℓ) hv
hu = σ(ℓ) ( ΣS p ), (c.8)
v∈N (u) {u} (1 + du )(1 + dv )

where W (ℓ) ∈ Rdℓ−1 ×dℓ is the learnable weight matrix with dℓ being the embedding dimension of layer ℓ and
σ(ℓ) is the activation function of ℓ-th layer. We recall that h(0) = X ∈ Rn×d is set to the initial node features.
Let X the original node features and denote by X ′ the perturbed adversarial features. Let’s consider a
node u ∈ V, we denote by hu its representation in the clean graph and h′u its representation in the attacked
graph , we note that the activation functions (σ(ℓ) )1≤ℓ≤ L are nonexpensive. Since we only consider node
feature based adversarial attacks, then by definition the original node u and its corresponding node in the
corrupted graph have the same neighborhood, we can therefore write

( L) ( L) (ℓ) (ℓ)
∥hu − h′ u ∥= ∥hu − h′ u ∥
(ℓ−1) (ℓ−1)
W (ℓ) hv W (ℓ) h′ v′
= ∥σ(ℓ) ( ΣS p ) − σ(ℓ) ( Σ p )∥
v∈N (u) {u} (1 + du )(1 + dv ) v′ ∈N (u) {u} (1 + du )(1 + dv′ )
S

(ℓ−1) (ℓ−1)
hv h ′ v′
≤ ∥W (ℓ)
∥∥ ΣS p − Σ p ∥
v∈N (u) {u} (1 + du )(1 + dv ) v′ ∈N (u) {u} (1 + du )(1 + dv′ )
S

(ℓ−1) (ℓ−1)
hv − h′ v
≤ ∥W (ℓ)
∥∥ ΣS p ∥
v∈N (u) {u} (1 + du )(1 + dv )
(ℓ−2)
1 W (ℓ−1) h j
= ∥W (ℓ)
∥∥ ΣS p [σ (ℓ−1)
( ΣS q )−
v∈N (u) {u} (1 + du )(1 + dv ) j∈N (v) {v} (1 + dv )(1 + d j )
(ℓ−2)
W (ℓ−1) h′ j
σ (ℓ−1)
( ΣS q )]∥
j∈N (v) {v} (1 + dv )(1 + d j )
(ℓ−1) (ℓ−1)
1 hj − h′ j
≤ ∥W (ℓ)
∥∥W (ℓ−1)
∥∥ ΣS p ΣS q ∥
v∈N (u) {u} (1 + du )(1 + dv ) j∈N (v) {v} (1 + dv )(1 + d j )

By iteration on the same process, we get the following result:

L
∥ h u − h ′ u ′ ∥ ≤ ∏ ∥W ( l ) ∥ 2 ∥
( L) (L
ΣS ΣS ...
l =1 v∈N (u) {u} j∈N (v) {v}

Xu − Xu′
ΣS p p ∥
z∈N (y) {y} (1 + du )(1 + dw )(1 + d j ) . . . (1 + dy ) (1 + d z )
L
≤ ∏∥W (l ) ∥∥ŵu ∥ϵ
l =1

with ŵu being the sum of normalized walks of length ( L − 1) starting from node u.
C.3 P R O O F O F T H E O R E M 5.4.1 AND 5.4.2 151

Let’s now consider the final output of the model which represents in the case of node classification the
individual output of each node.
 .. 
.
∥ f ( A, X ) − f ( A, X ′ )∥= ∥
 ( L) 
h − h ′ ( L)  ∥
 u u 
..
.
Based on the previous formulation, we have the following results:

L
∥ f ( A, X ) − f ( A, X ′ )∥1 ≤ ϵ ∏∥W (l ) ∥1 ∑ ŵu
l =1 u∈V

L
∥ f ( A, X ) − f ( A, X ′ )∥∞ = ϵ ∏∥W (l ) ∥∞ max ŵu
u∈V
l =1

Therefore, we computed the Lipschitz constant of our considered model f . To link this constant to our
robustness definition in Equation 5.1, we use the Markov inequality as follows

Advϵ [ f ] = P(G,X )∼DG ,X [( G̃, X̃ ) ∈ Bα,β ( G, X, ϵ) s.t. dY ( f ( G̃, X̃ ), f ( G, X )) > σ ]


α,β

1
≤ E (G,X )∼DG ,X , ∥ f ( A, X ) − f ( A, X ′ )∥∞
 
σ (G′ ,X ′ )∈ B ((G,X ),ϵ)
α,β
!
L
1
σ ∏
(l )
≤ ∥W ∥op ϵ max | ŵu |1 .
u∈V
l =1

Thus, the classifier f is ((d0,1 , ϵ), (d∞ , γ))–robust robust with


L
γ= ∏∥W (i) ∥ϵ max
u∈V
ŵu /σ.
i =1

And we also have that the classifier f is ((d0,1 , ϵ), (d1 , γ))–robust with
L
γ= ∏∥W (i) ∥ϵ( ∑ ŵu )/σ.
i =1 u∈V

Theorem c.3.1. Let f : (G , X ) → Y denote a graph function composed of L GCN layers, where W (i)
denotes the weight matrix of the i-th layer. Further, let d1,0 be a graph distance. For structural attacks with
a budget ϵ, the function f is ((d1,0 , ϵ), (d2 , γ))–robust with
L L
γ= ∏∥W (i) ∥∥X ∥ϵ(1 + L ∏∥W (i) ∥)/σ.
i =1 i =1

Proof. Following the same analogy, let’s consider the case of structural perturbations. We only consider
that the activation functions (σ(ℓ) )1≤ℓ≤ L are nonexpensive. In this setting, the graph shift operator (the
normalized adjacency matrix in our work) is the one to be edited and the input features are shared between
(i )
the two graphs. Let à the clean adjacency and h(i) its hidden represented in the i-th layer and Ã′ , h′ ,
with the attacked/perturbed one.
152 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

Let’s first consider the following

(l ) ( l −1)
∥h′ ∥ = ∥σ(l ) ( Ã2 h′ W (l ) )∥
( l −1)
≤ ∥ Ã2 ∥∥W (l ) ∥∥h′ ∥
( l −1)
≤ ∥W (l ) ∥∥h′ ∥.
( L)
By recursion, we have: ∥h′ ∥ ≤ ∏iL=1 ∥W (i) ∥∥ X ∥.
From another side, we have the following

(l ) (l ) (l )
∥σ(l ) ( Ãh1 W (l ) ) − σ(l ) ( Ã′ h′ 2 W (l ) )∥2 ≤ ∥ Ãh(l ) W (l ) − Ã′ h′ 2 W (l ) ∥2
(l ) (l ) (l )
≤ ∥W (l ) ∥∥ Ãh(l ) − Ã′ h′ + Ãh′ − Ãh′ ∥2
(l ) (l )
≤ ∥W (l ) ∥∥ Ã(h(l ) − h′ ) − h′ ( Ã − Ã)∥2
(l ) (l )
≤ ∥W (l ) ∥(∥ Ã∥∥(h(l ) − h′ )∥+∥h′ ∥∥( Ã − Ã′ )∥2 )
(l ) (l )
≤ ∥W (l ) ∥∥(h(l ) − h′ )∥+∥W (l ) ∥∥h′ ∥ϵ.

By recursion, we get the following (By considering ∀i > 1 : ∥W (i) ∥ ≥ 1)

L L
∥Φ( A1 , X ) − Φ( A2 , X )∥2 ≤ ∏∥W (i) ∥∥ X ∥ϵ + L(∏∥W (i) ∥)2 ∥ X ∥ϵ
i =1 i =1
L L
≤ ∏∥W (i) ∥∥ X ∥ϵ(1 + L ∏∥W (i) ∥).
i =1 i =1

Similarly to the previous proof, we deduce that the classifier f is ((d1,0 , ϵ), (d2 , γ))–robust where

γ = ∏iL=1 ∥W (i) ∥∥ X ∥ϵ(1 + L ∏iL=1 ∥W (i) ∥)/σ.

C.4 PROOF OF THEOREM 5.4.3

Proof. Let’s now consider a graph-function f that is based on L GIN-layers (with a parameter ζ = 0, that is
usually denoted by ϵ in the literature). The GIN message-passing propagation process can be written for a
node u as:

(ℓ+1) (ℓ) (ℓ)


hu = T (ℓ+1) ((1 + ζ )hu + Σ hv )
v∈N (u)

where T denotes a Neural Networks (a MLP) for example and ζ denotes the parameter of the GIN. We
recall that h(0) = X ∈ Rn×d is set to the initial node features.
Let X the original node features and X̃ the perturbed features. Let’s consider a node u ∈ V, we denote
by hu its representation in the clean graph and h′u its representation in the attacked graph , we note that
we consider the activation functions to be nonexpensive. Since we only consider node feature based
C.4 P R O O F O F T H E O R E M 5.4.3 153

adversarial attacks, then by definition the original node u and its corresponding corrupted node have the
same neighborhood, we can therefore write

( L) ( L) (ℓ−1) (ℓ−1) (ℓ) (ℓ)


∥hu − h′ u ∥ = ∥ T (ℓ) ((1 + ζ )hu + Σ hv ) − T (ℓ) ((1 + ζ )h′ u + Σ h′ v )∥
v∈N (u) v∈N (u)
(ℓ−1) (ℓ−1) (ℓ−1) (ℓ−1)
≤ ∥W (ℓ) ∥∥(1 + ζ )(hu − h′ u )+ Σ ( hv − h′ v )∥
v∈N (u)

We assume that the input feature space H0 is bounded, thus each hidden space Hi of the iterative
process of message passing is bounded and let B = max Bℓ be its global maximum bound. From the
ℓ≤ L
previous result, we have

(ℓ) (ℓ) (ℓ−1) (ℓ−1)


∥hu − h′ u ∥≤ ∥W (ℓ) ∥∥(1 + ζ )(hu − h′ u ) + B × deg(u)∥

By iteration over the previous process, and by considering ∀i > 1 : ∥W (i) ∥ ≥ 1, we can write

L L L
∥hu − h′ u ∥ ≤ B × deg(u) × ∏∥W (l ) ∥ ∑ (1 + ζ )l + ∏∥W (l ) ∥(1 + ζ ) L ∥hu − h′ u ∥
( L) ( L) (0) (0)

l =1 l =1 l =1
L L L
≤ B × deg(u) × ∏∥W (l ) ∥ ∑ (1 + ζ )l + ∏∥W (l ) ∥(1 + ζ ) L ϵ
l =1 l =1 l =1
L L
≤ ∏∥W (l ) ∥[ B × deg(u) ∑ (1 + ζ )l + (1 + ζ ) L ϵ]
l =1 l =1

Since we consider the GIN-parameter ζ ≈ 0, we can deduce from the previous result that
L
∥≤ ∏∥W (l ) ∥[ B × L × deg(u) + ϵ]
(ℓ+1) (ℓ+1)
∥ hu − h ′ u′
l =1

Let’s now consider the final output of the model which represents in the case of node classification the
individual output of each node.
 .. 
.
∥ f ( A, X ) − f ( A, X ′ )∥= ∥
 ( L) 
′ ( L) 
 hu − h u  ∥
..
.

Based on the previous formulation, we have the following results

L
∥ f ( A, X ) − f ( A, X ′ )∥∞ = ∏∥W (l ) ∥∞ [ B × L × max deg(u) + ϵ]
u∈V
l =1

Similar to the previous proof of Theorem 5.4.1, we connect the Lipschitz constant of our considered
model f to the robustness definition in Equation 5.1 through Markov inequality and we get the following

L
γ= ∏∥W (l) ∥∞ [ B × L × max
u∈V
deg(u) + ϵ]/σ.
l =1
154 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

C.5 G E N E R A L I Z I N G TO A N Y G R A P H N E U R A L N E T WO R K

While this work focuses on GCNs, it can be easily extended to other GNN architecture. Given the iterative
nature of GNNs, they can be viewed as a composition of multiple continuously differentiable functions,
f i : Hi → Hi+1 , where Hi represents the input space of the i − th function. The Lipschitz constant, as
shown in Zhao et al., 2021 is given by C = ∏i ∥▽ f i ∥. The functions { f i }i can be grouped into two categories
: (i) The set L of functions f i that are linear with a corresponding parameter Wi . (ii) The set of functions
that are not linear. We can hence decompose C ′ according to the nature of the function as follows
  
C = ∏ ∥▽ f i ∥

∏ ∥Wi ∥ .
fi ∈
/L f i ∈L

We can assume the input feature space H0 to be bounded; this a realistic assumption. Thus, each
hidden space Hi is also bounded, and since {▽ f i } f i ∈/ L are continuous functions on bounded spaces, there
exists an upper bound Ci for ∥▽ f i ∥. Therefore, there exists C > 0, such that,

C′ ≤ C ∏ ∥Wi ∥.
f i ∈L

We return to the previous case by simply taking ϵ′ = ϵC/C ′ . Furthermore, following the assumption on the
input feature space being bounded, it is possible to derive an upper bound on the GNN’s robustness when
subject to both structural and feature-based attacks simultaneously (See Appendix c.6).

C.6 V U L N E R A B I L I T Y U P P E R - B O U N D W H E N D E A L I N G W I T H B OT H S T RU C T U R A L A N D N O D E
F E AT U R E S AT TA C K S

In this section, we derive the upper-bound for a GCN’s vulnerability Advα,β,ϵ [ f ] when dealing with both
structural and feature-based attacks simultaneously.

1
Advα,β,ϵ [ f ] ≤ E ( G1 ,X1 )∼DG ,X , [∥ f ( A1 , X1 ) − f ( A2 , X2 )∥2 ]
σ (G
2 ,X2 )∈ Bα,β (( G1 ,X1 ) ,ϵ )

1
≤ E ( G1 ,X1 )∼DG ,X , [∥ f ( A1 , X1 ) − f ( A1 , X2 ) + f ( A1 , X2 ) − f ( A2 , X2 )∥2 ]
σ (G
2 ,X2 )∈ Bα,β (( G1 ,X1 ) ,ϵ )

1
≤ E ( G1 ,X1 )∼DG ,X , [∥ f ( A1 , X1 ) − f ( A1 , X2 )∥2 +∥ f ( A1 , X2 ) − f ( A2 , X2 )∥2 ]
σ (G
2 ,X2 )∈ Bα,β (( G1 ,X1 ) ,ϵ )

1
≤ E ( G1 ,X1 )∼DG ,X , [∥ f ( A1 , X1 ) − f ( A1 , X2 )∥2 ] +
σ (G
2 ,X2 )∈ Bα,β (( G1 ,X1 ) ,ϵ )

E ( G1 ,X1 )∼DG ,X , [∥ f ( A1 , X2 ) − f ( A2 , X2 )∥2 ] .


( G2 ,X2 )∈ Bα,β (( G1 ,X1 ),ϵ)

We already established an upper-bound in Appendix c.3 for each term of the previous inequality.
Therefore, we can have a new upper-bound for Advα,β,ϵ [ f ]

! !
L L L
1 1
Advα,β,ϵ [ f ] ≤
σ ∏ ∥W (i )
∥ϵ( ∑ wˆu ) 2
ϵ+
σ ∏ ∥W (i )
∥∥ X ∥ϵ(1 + L ∏∥W ∥) .(i )
i =1 u∈V i =1 i =1
C.7 T I M E A N D C O M P L E X I T Y A N A LY S I S 155

Table c.1 – Attacked classification accuracy (± standard deviation) of the models on different benchmark node
classification datasets after both structural attacks and node feature attacks application.
Dataset GCN GCN-Jaccard RGCN GNN-SVD GNNGuard ParsevalR GCORN
Cora 76.7 ± 1.2 76.7 ± 0.4 78.1 ± 0.6 65.9 ± 4.0 61.0 ± 0.4 77.6 ± 1.1 79.8 ± 0.7
CiteSeer 68.7 ± 0.3 71.5 ± 0.3 68.9 ± 0.8 68.3 ± 0.5 65.0 ± 0.3 70.9 ± 0.6 74.2 ± 0.3
PubMed 47.0 ± 0.4 49.4 ± 0.8 51.2 ± 0.4 47.7 ± 0.5 40.3 ± 0.3 49.7 ± 1.1 53.9 ± 0.8
CoraML 63.3 ± 0.0 32.2 ± 1.5 59.7 ± 1.4 53.8 ± 0.8 32.2 ± 1.5 65.4 ± 0.2 77.4 ± 1.1

Following the assumption on the input feature space being bounded, i.e. ∃ B > 0, ∀ X ∥ X ∥ ≤ B, the
inquality become

! !
L L
1
Advα,β,ϵ [ f ] ≤
σ ∏ ∥W (i )
∥ ( ∑ wˆu ) + B(1 + L ∏∥W ∥) ϵ.
2 (i )
i =1 u∈V i =1

Thus, the classifier f is d0,1 -(ϵ, γ) robust where


! !
L L
γ= ∏ ∥W ( i ) ∥ ( ∑ wˆu )2 + B(1 + L ∏∥W (i) ∥) ϵ/σ.
i =1 u∈V i =1

C.6.1 Experimental Validation

We assessed the efficacy of our proposed method against both Structural and Node-feature-based
adversarial attacks simultaneously. This evaluation involved a combined evasion attack approach, wherein
two attacks were integrated. Initially, we generated the perturbed adjacency matrix by employing a surrogate
GCN model along with the "Mettack" (utilizing the ‘Meta-Self’ strategy) (Zügner and Günnemann, 2019).
Subsequently, we perturbed the node features using a random attack strategy involving the injection of
Gaussian noise N (0, I) into the features. This perturbation was controlled by a scaling parameter ψ, which
we identified as effective in prior experiments. For the structural perturbations, we set the perturbation
budget to 0.1E, while for the node features, ψ was set to 5.0. We compared our method to the same
considered defense benchmarks that were used in Section 5.6.

C.7 T I M E A N D C O M P L E X I T Y A N A LY S I S

Drawing from the original paper (Björck and Bowie, 1971), it has been established that the Bjorck
orthonormalization method will always converge, provided that the condition ∥W T W − I ∥2 < 1 is met. In
this section, we will begin by examining the connection between the selected order, the number of iterations
and the resulting performance and then present an empirical time analysis of the proposed GCORN on
various datasets.

C.7.1 On the Effect Of Order/Iterations

As per the projection equation 5.4, it appears that a higher order/iteration generally results in a more
precise projection of the weight matrix. In our study, we gauge the accuracy of the approximation by
assessing the model’s performance in the absence of adversarial attacks, as well as under attack. Although
156 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

our experiments demonstrate that employing the first order with a restricted number of iterations can produce
satisfactory outcomes, we believe that an hyper-parameter analysis should be conducted. Consequently,
we evaluate the influence of modifying the order (while maintaining a constant number of iterations) on the
resulting accuracy in Table c.2.

Table c.2 – Performance of GCN and our proposed GCORN model, for different used approximation orders, on the
Cora dataset.

GCN GCORN(1 ord) GCORN(2 ord) GCORN(3 ord)

Training Time (in s) 2.8 ± 0.01 4.8 ± 0.07 8.7 ± 0.07 10.9 ± 0.08
Accuracy w/o attack 79.2 ± 1.6 78.8 ± 1.3 79.8 ± 0.9 80.8 ± 1.1
Accuracy w. attack 68.4 ± 1.9 77.1 ± 2.1 78.3 ± 1.1 78.6 ± 0.4

Table c.2 shows both clean and attacked accuracy of the GCN and our proposed GCORN and the
corresponding standard deviations for different orders. We observe that using higher order approximations
generally leads to enhanced accuracy and robustness, albeit at the expense of increased running time.
Therefore, the choice of order hyperparameter employed in the approximation must be carefully tuned
for each application to find the best trade-off between robustness,accuracy and training time. In our
experiments, we found that an approximation of order 1 yielded adequate and satisfactory results.

C.7.2 On Training Time

In line with the above, we also investigate the training time of our model compared to that of a standard
GCN. To conduct this experiment, we held the order and number of iterations constant and utilized identical
architectures and hyperparameters for both models (see Appendix c.12).

Table c.3 – Mean training time analysis (in s) of a our GCORN in comparison to the other benchmarks.

Dataset GCN GCN-K AIRGNN RGCN GCORN

Cora 2.8 1.8 2.6 3.2 4.8


CiteSeer 2.4 5.8 2.9 2.4 4.6
PubMed 5.9 8.9 7.4 14.5 7.3
CS 6.1 12.1 12.4 13.8 15.5
Ogbn-Arxiv 77.8 185.8 68.1 161.6 78.4

The mean training time are presented in Table c.3, indicating a trade-off between the improved robustness
provided by the GCORN model and its slightly longer training time compared to the GCN model and other
available methods.

C.8 M O R E D E TA I L S A B O U T T H E E S T I M AT I O N O F O U R R O B U S T N E S S M E A S U R E

We begin by providing the pseudocode of the algorithm used to estimate our robustness measure.
C.8 M O R E D E TA I L S A B O U T T H E E S T I M AT I O N O F O U R R O B U S T N E S S M E A S U R E 157

α,β
Algorithm 6: Estimation of Advϵ [ f ].
1 Inputs: Sphere Radius : ϵ > 0, Number of Samples Lmax , Number of Input Graphs |D|;
2 Initialize Adv = 0;
3 foreach [ Gi , Xi ] ∈ D do
4 Initialize Advi = 0;
5 foreach l = 1, . . . , Lmax do
6 1. Sample a distance r ∈ [0, ϵ] from the prior distribution pϵ (see Appendix c.9);
7 2. Uniformly sample Zl ∈ Rn×K from Sr (see Appendix c.8);
8 3. Choose X̃l = Xi + Zl ;
9 4. Update
10 Advi ← Advi + 1{dY ( f ( G̃l , X̃l ), f ( G, X )) > σ };
11 Adv = Adv + Advi /Lmax ;
12 Return Adv/|D| ;
We now detail how to uniformly sample from the set Sr defined by,
Sr = { Z ∈ R n × D | max ∥ Zi ∥ p = r },
1≤ i ≤ n

where r ∈ [0, ϵ]. The set Sr can also be defined in the following way,

Z ∈ Sr ⇔ max ∥ Zi ∥ p = r
1≤ i ≤ n
(
∃i0 ∈ {1, . . . , n} ∥ Zi0 ∥ p = r,

∀i ∈ {1, . . . , n} ∥ Zi ∥ p ≤ r.
Since the position of i0 within {1, . . . , n} do not matter in the previous equivalence, we can uniformly
sample i0 ∈ {1, . . . , n}, such that ri0 = ∥ Zi0 ∥ p = r and that the other index i ∈ {1, . . . , n} \ {i0 } should
satisfy ri = ∥ Zi ∥ p ≤ r, i.e., Zi ∈ BK (r ) = { x ∈ RK |∥ x ∥ p ≤ r }. We will use another time Stratified Sampling
where we start by sampling {ri }i̸=i0 within [0, r ]. Using the Lemma 5.5.1, we can directly sample {ri }i̸=i0
from the probability distribution

1  r i  K −1
∀i ∈ {1, . . . , n} \ {i0 }, p(ri ) = K .
r r
Once we know the values of {ri }i , the problem boils down to sampling n vectors { Zi }1≤i≤n from RK ,
such that

∀i ∈ {1, . . . , n}, ∥ Zi ∥ p = ri . (c.9)


For that, we have to consider the two cases p < ∞ and p = ∞

C.8.1 Case Where p < ∞

In this case, Equation c.9 can be written as follows

K
∑ |Zi,j | p = ri ,
p
∀i ∈ {1, . . . , n}, ∥ Zi ∥ p = ri ⇔ ∀i ∈ {1, . . . , n},
j =1
K p
| Zi,j |

⇔ ∀i ∈ {1, . . . , n}, ∑ ri
= 1. (c.10)
j =1
158 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

p1 p2 p4

0 O1 O2 O3 O4 1

Figure c.1 – A toy example on how to randomly partition the set [0, 1] into K = 4 parts such as the sum of the
parts lengths is 1. We first uniformly sample 3 elements from [0, 1] and reorder them [ p1 , p2 , p3 ]. And
subsequently, we consider O1 = [0, p1 ], O2 = [ p4 , p3 ], O3 = [ p4 , p3 ], O4 = [1, p4 ].

To satisfy Equation (c.10), we randomly uniformly partition the [0, 1] into K parts. To do so, for each i ∈
(i ) (i )
{1, . . . , n}, we randomly sample D − 1 element from uniform distribution U [0, 1]; Let’s denote [ p1 , . . . , pK−1 ]
the sorted list of sampled elements. We directly choose
(i )


 1 − p j −1 if j = K,

(i )
Oi,j = pj if j = 0,

 (i )
 (i )
p j − p j−1 otherwise .

The matrix O = Oi,j 1≤i ≤n,1≤ j≤ D
satisfy

K
∀i ∈ {1, . . . , n}, ∑ Oi,j = 1.
j =1

Figure c.1 gives insight into the previously introduced concept of randomly partitioning in the case of the
set [0, 1] into K = 4.
Based on the previous elements, we can choose Z such that

| Zi,j | p
 
1/p
∀i, j, = Oi,j ⇔ ∀i, j, | Zi,j | = ri × Oi,j . (c.11)
ri

Equation c.11 is invariant with respect to the sign of Zi,j , hence, we use the following
1/p
∀i, j, Zi,j = ui,j × ri × Oi,j ,

where ui,j is sampled from the discrete uniform distribution U{−1,1} .


We recall that for p = 1, the distance d0,1 ([ G, X ], [ G̃, X̃ ]) = maxi∈{1,...,n} ∥ Xi − X̃i ∥ p defined in Equation
∥ AX − BX ∥∞
5.6 matches with the infinity matrix distance d∞ : ( A, B) 7→ ∑ X ̸=0 ∥ X ∥∞
as shown in Nagisa, 2022.
α,β
Therefore, from the Theorem 5.4.1, the upper-bound of the expected vulnerability Advϵ [ f ] is γ =
∏iL=1 ∥W (i) ∥∞ ϵŵG /σ.

C.8.2 Case Where p = ∞

In this case, Equation c.9 can be written as follows

∀i ∈ {1, . . . , n}, ∥ Zi ∥∞ = ri ⇔ ∀i ∈ {1, . . . , n}, max | Zi,j | = ri


1≤ j ≤ K
(
∃ j0 ∈ {1, . . . , K } | Zi,j0 | = ri ,
⇔ ∀i ∈ {1, . . . , n}, (c.12)
∀ j ∈ {1, . . . , K } | Zi,j0 | ≤ ri .
C.9 P R O O F O F L E M M A 5.5.1 159

For each i ∈ {1, . . . , n}, we randomly select j0 ∈ {1, . . . , K } to satisfy | Zi,j0 | = ri . For the other j ̸= j0 , we
uniformly sample the value of | Zi,j | from U[0,ri ] . These equations are invariant with respect to the sign of
Zi,j , therefore, we choose,
∀i, j, Zi,j = ui,j × | Zi,j |,
where ui,j is sampled from the discrete uniform distribution U{−1,1} .

C.9 PROOF OF LEMMA 5.5.1

Proof. Using the following equivalence


∃ Z ∈ Rn,K max ∥ Zi ∥ p = r ⇔ ∃ T ∈ RK , ∥ T ∥ p = r.
1≤ i ≤ n

We deduce that the density pϵ do not depend on n, and therefore we can set n = 1 for this proof. Therefore,
n o
Bϵ = Z ∈ RK |∥ Z ∥ p ≤ ϵ .

We start by calculating the volume of the K −sphere of radius r for any real finite-dimensional space RK
where K > 2 is an integer.
SK (r ) = { x ∈ RK |∥ x ∥ p = r }.
Let’s denote the volume of SK (r ) by V K (r ).
Using Fubini Theorem, we have

Z r  
V K (r ) = V K−1 (r p − | x | p )1/p dx
−r
Z 1  
=r V K−1 r (1 − | x | p )1/p dx
−1
Z 1  
1/p
=r r K −1 V K −1 (1 − | x | p ) dx
−1
Z 1  
= rK V K−1 (1 − | x | p )1/p dx
−1
K K
= r V (1).
(c.13)
Thus, the surface L p ball in RK of radius r is

d V K (r )
S K (r ) = (c.14)
dr
= Kr K−1 V K (1). (c.15)

Consequently the density distribution of R( p) could be written as


p
S (r )
pϵ (r ) = Kp 1 {0 ≤ r ≤ ϵ }
VK (ϵ)
p
1 d V K (r )
p = 1 {0 ≤ r ≤ ϵ }
VK (ϵ) dr
1  r  K −1
=a 1 {0 ≤ r ≤ ϵ }. (c.16)
ϵ ϵ
We note that the previous quantity doesn’t depend on p.
160 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

C.10 A D D I T I O N A L R E S U LT S

C.10.1 Node Classification

To further understand the relationship between robustness and input sensitivity, from which we derived
the upper bound for GNNs and specifically GCNs, we empirically validate the trade-off between robustness
and input perturbation. We hence compare the difference in the output of each model when subjected
to random perturbations with varying attack budgets. The results of this study are presented in Figure
c.2. The figure displays the results for both Cora and CiteSeer datasets in both absolute (left figure) and
log-scale (right figure) format. The results demonstrate that the GCN is highly sensitive to perturbations in
comparison to the proposed approach in terms of the difference in output. These findings provide validation
for the stability of GCORN against input perturbations, and by extension, adversarial attacks.
We additionally considered two defense techniques benchmarks: (1) ElasticGNN (Liu et al., 2021c)
which proposes a novel and general message passing scheme into GNNs to enhance the local smoothness
adaptivity of GNNs via ℓ1 -based graph smoothing and (2) EvenNet (Lei et al., 2022) which proposes a
spectral GNN corresponding to even-polynomial graph filter with the main idea being that ignoring odd-hop
neighbors improves the underlying robustness. For this additional evaluation, we consider similar attacks
techniques as the one from Table 5.1.

Table c.4 – Attacked classification accuracy (± standard deviation) of the models on different benchmark node
classification dataset after the attack application.
Attack Dataset ElasticGNN EvenNet GCORN
Cora 72.5 ± 2.4 69.7 ± 1.8 77.1 ± 1.8
Random
CiteSeer 58.9 ± 2.3 48.6 ± 1.6 67.8 ± 1.4
(ψ = 0.5)
PubMed 70.5 ± 0.6 68.8 ± 1.6 73.1 ± 1.1
Cora 54.4 ± 2.4 48.9 ± 2.1 57.6 ± 1.9
Random
CiteSeer 43.8 ± 2.0 39.7 ± 4.2 57.3 ± 1.7
(ψ = 0.5)
PubMed 56.8 ± 2.1 58.5 ± 3.2 65.8 ± 1.4
Cora 64.7 ± 1.6 57.6 ± 3.6 71.1 ± 1.4
PGD CiteSeer 50.8 ± 3.2 37.1 ± 3.3 65.6 ± 1.4
PubMed 64.6 ± 1.5 61.0 ± 4.1 72.3 ± 1.3

C.10.2 Graph Classification

In this section, we present the result of our GCORN on the graph classification task using both empirical
and our proposed probabilistic evaluation. Details about the experimental setting for this task are provided
in Appendix c.12. Note that for the random attack, we used ψ = 0.5 and for the gradient-based, we used a
budget δ = 0.2. Table c.5 reports the average accuracy and the corresponding standard deviation for both
clean and attacked accuracy.
C.11 E X P E R I M E N TA L R E S U LT S O N G I N 161

Table c.5 – Classification accuracy (± standard deviation) of the models on different benchmark graph classification
dataset before (clean) and after the attack. The higher the accuracy (in %) the better.
PROTEINS D&D NCI1
Attack
GCN GCORN GCN GCORN GCN GCORN

Clean 73.4 ± 2.8 74.1 ± 1.9 75.8 ± 3.6 76.4 ± 4.1 75.7 ± 2.2 74.8 ± 1.7
Random 66.7 ± 2.5 71.7 ± 2.8 70.9 ± 3.2 75.1 ± 4.8 67.1 ± 2.6 71.8 ± 2.1
= PGD 56.7 ± 2.8 65.4 ± 3.1 61.8 ± 4.1 68.6 ± 4.3 54.9 ± 2.9 62.4 ± 2.1

C.11 E X P E R I M E N TA L R E S U LT S O N G I N

To showcase our method’s versatility beyond the considered GCN framework and confirm the GIN-
related results outlined in Theorem 5.4.3, we conducted similar experiments to those presented in Table
5.1. However, this time, we employed GIN as the message-passing scheme. In this perspective, we
used the same previously considered node features attacks, notably: (i) The baseline random attack
injecting Gaussian noise N (0, I) to the features with a scaling parameter ψ controlling the attack budget;
(ii) The white-box Proximal Gradient Descent (Xu et al., 2019a), which is a gradient-based approach to the
adversarial optimization task with a chosen attack budget of 15%.

Table c.6 – Attacked classification accuracy (± standard deviation) of the models on different benchmark node
classification dataset using Graph Isomorphism Network after the attack application.
Model Cora CiteSeer PubMed
GIN 67.8 ± 1.3 50.1 ± 0.3 72.8 ± 0.4
Random (ψ = 0.5)
GIORN 71.8 ± 0.8 60.8 ± 0.5 74.3 ± 0.7
GIN 54.5 ± 0.9 45.1 ± 0.4 67.3 ± 0.6
Random (ψ = 1.0)
GIORN 66.7 ± 1.2 54.4 ± 0.8 68.9 ± 1.1
GIN 61.4 ± 1.4 44.1 ± 0.5 70.6 ± 0.4
PGD
GIORN 69.5 ± 0.6 58.3 ± 0.6 73.1 ± 1.6

C.12 D ATA S E T S A N D I M P L E M E N TAT I O N D E TA I L S

For all the used models, the same number of layers, hyperparameters, and activation functions were
used. The models were trained using the cross-entropy loss function with the Adam optimizer, the number
of epochs and learning rate were kept similar for the different approaches across all experiments.

C.12.1 Node Classification

Characteristics and information about the datasets utilized in the node classification part of the study are
presented in Table c.8. As outlined in the main paper, we conduct experiments on a set of citation networks,
including Cora, CiteSeer, and PubMed (Sen et al., 2008), as well as the Amazon Co-author network of
authors from the Computer Science (CS) domain (Shchur et al., 2018b). For Cora, CiteSeer, and PubMed,
we adhere to the train/valid/test splits provided by the datasets. For the CS dataset, we adopt the same
162 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

Table c.7 – Statistics of the node classification datasets used in our experiments.

Dataset #Features #Nodes #Edges #Classes

Cora 1433 2708 5208 7


CiteSeer 3703 3327 4552 6
PubMed 500 19717 44338 3
CS 6805 18333 81894 15
OGBN-Arxiv 128 31971 71669 40

Table c.8 – Statistics of the graph classification datasets used in our experiments.

Dataset #Graphs #Nodes #Edges #Classes

DD 1178 284.32 715.66 2


NCI1 4110 29.87 32.30 2
PROTEINS 1113 39.06 72.82 2

methodology as used in Yang, Cohen, and Salakhudinov (2016), by randomly selecting 20 nodes from
each class to form the training set and 500/1000 nodes from the remaining for the validation and test sets.
In all of the experiments, the models employed a 2-layer convolutional architecture (consisting of two
iterations of message passing and updating) stacked with a Multi-Layer Perception (MLP) as a readout.
The intent was to compare the models in an iso-architectural setting, to ensure a fair evaluation of their
robustness. All experiments were conducted using the Adam optimizer (Kingma and Ba, 2015b) and the
same hyperparameters, including a learning rate of 1e-2, 300 epochs, and a hidden feature dimension
of 16. For the OGB dataset, we used 512 as a hidden feature dimension. To account for the impact of
random initialization, each experiment was repeated 10 times, and the mean and standard deviation of the
results were reported. We finally note that for the AIRGNN, we set K = 2 so as to have the same number
of propagation as the other benchmarks. For our proposed GCORN, we tuned the iteration numbers for
the different datasets and this was also done for the Parseval Regularization (ParselR) to find the best
regularization parameter.

C.12.2 Graph Classification

We additionally emperically evaluated our proposed method GCORN using both probabilistic evaluation
and the classical experimental evaluation. We used benchmark datasets derived from bioinformatics and
chemoinformatics (PROTEINS, NCI1, D&D) (Morris et al., 2020a). The framework proposed by Errica et al.
(2020) was used to evaluate the performance of the models on this task. We therefore performed a 10-fold
cross-validation using the same folds as provided by the paper to obtain an estimate of the generalization
performance of each method.
Similar to node classification, we employed a 2-layer convolutional architecture (consisting of two
iterations of message passing and updating) stacked with a Multi-Layer Perception (MLP) as a readout
using the Adam Optimizer.
C.12 D ATA S E T S A N D I M P L E M E N TAT I O N D E TA I L S 163

C.12.3 Implementation Details

Our implementation is available in the supplementary materials (and will be publicly available afterwards).
It is built using the open-source library PyTorch Geometric (PyG) under the MIT license (Fey and Lenssen,
2019). We leveraged the publicly available implementation of the different benchmarks from their available
repositories : From GCN-k 1 , for AIRGNN 2 , for GNNGuard 3 and for RGCN we used the implementation
from the DeepRobust package. Note that we additionally utilized the PyTorch DeepRobust package 4 to
implement the adversarial attacks used in this study. The experiments have been run on both a NVIDIA
A100 GPU and a RTX A6000 GPU.

C.12.4 Implementation Details of our Empirical Robustness Estimation

We considered the input distance defined in Equation 5.6 with p = 2, we fixed the radius ϵ = 10 and the
number of sampling per graph at Lmax = 100. For each graph G, the output distance dY has been re-scaled

to the interval [0, 1] by normalization using a factor of 2 NG where NG is the number of nodes in the graph
G.

1. [Link]
2. [Link]
3. [Link]
4. [Link]
164 A P P E N D I X : A DV E R S A R I A L R O B U S T N E S S O F G N N S

Figure c.2 – Difference (in average and standard deviation) in output for the GCN and our GCORN when subject to
random perturbations with different attack budgets for both (a) Cora and (b) CiteSeer. The right-hand
side plots are on the log-scale.
A P P E N D I X : A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M
FIELDS
d

165
166 A P P E N D I X : A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

D.1 PROOF OF LEMMA 6.3.1

Lemma 6.3.1 By solving the system of equations in (6.3), we can get the optimal distribution Q∗ as
follows: n h  io
ea ) ∝ exp E−a log P Y
∀ a ∈ V CRF , Q(Y e|Y, V CRF , ECRF .

e|Y, V CRF , ECRF ) which is exactly the unknown


Proof. The KL-divergence includes the true posterior P(Y
value. We can rewrite the KL-divergence as:

Q (Y
e)
Z
KL ( Q| P) = Q (Y
e) log dY
e
e|Y, V CRF , ECRF )
P (Y
Z e) P(Y, V CRF , ECRF )
Q (Y
= Q (Y
e) log dY
e
e Y, V CRF , ECRF )
P(Y,
!
Q (Y
e)
Z
CRF CRF
= Q (Y
e) log P(Y, V ,E ) + log dYe
e Y, V CRF , ECRF )
P(Y,
CRF , ECRF )
e) log P(Y, Y, V
Z Z e
= log P(Y, V CRF , ECRF ) Q(Y e − Q (Y
e ) dY dY.
e
Q (Y
e)
R
Since Q (Y
e ) dY
e = 1, we conclude that:

CRF , ECRF )
e) log P(Y, Y, V
Z e
CRF CRF
KL ( Q| P) = log P(Y, V ,E )− Q (Y dY.
e
Q (Y
e)

We are minimizing the KL−divergence over Q, therefore the term log P(Y, V CRF , ECRF ) can be ignored.
The second term is the This is the negative ELBO. We know that the KL−divergence is not negative. Thus,
log P(Y, V CRF , ECRF ) ≥ ELBO( Q) justifying the name Evidence lower bound (ELBO).
Therefore, the main objective is to optimize the ELBO in the mean field variational inference, i.e., choose
the variational factors that maximize ELBO:

Z e Y, V CRF , ECRF )
P(Y, h i h i
ELBO( Q) = Q (Y
e) log e = EQ log P(Y,
dY e Y, V CRF , ECRF ) − EQ log Q(Y
e) .
Q (Y
e)
(d.1)
We will employ coordinate ascent inference, where we iteratively optimize each variational distribution
while keeping the others constant.
We assume that the set of CRF nodes is finite, i.e., V CRF < ∞, which is a realistic assumption if we
consider the set of all GNN inputs used during inference. If V CRF = m, we can order the elements
V CRF in a a specific order i = 1, . . . , m. Thus, using the chain rule, we decompose the probability
e Y, V CRF , ECRF ) as follows
P(Y,

e Y, V CRF , ECRF ) = P(Y


P(Y, e1:m , Y1:m , V CRF , ECRF )
m
= P(Y1:m , V CRF , ECRF ) ∏ P(Y
ei |Y1:(i−1) , V CRF , ECRF ).
i =1

Using the independence of the mean field approximation, we also have


h i m h i
EQ log Q(Y
e) = ∑ EQ j
log Q(Y
ej ) .
i =1
D.2 A S Y M P T O T I C B E H AV I O R O F C R F N E I G H B O R H O O D S I Z E 167

Now, we have the expression of the two terms appearing in ELBO in (d.1):
m h i h i
ELBO( Q) = log P(Y1:m , V CRF , ECRF ) + ∑ EQ log P(Y
ei |Y1:(i−1) , V CRF , ECRF ) − EQ log Q(Y
j
ei ) .
i =1

The above decomposition is valid for any ordering of the GNN inputs. Thus, for a fixed GNN input a ∈ V CRF ,
if we consider a as the last variable m of the list, we can consider the ELBO as a function of Q(Y
ea ) = Q(Y
em )

ELBO( Q(Y
ea )) = ELBO( Q(Y ek ))
h i h i
= EQ log P(Y em |Y1:(m−1) , V CRF , ECRF ) − EQ log Q(Y
m
em ) + const
Z h i Z
= Q (Y em )EQ log P ( em |Y̸=m , V CRF , ECRF ) dY
Y em − Q(Y em )log Q(Yem )dY em
̸=m
Z h i Z
= Q (Y ea )EQ log P(Yea |Y̸=a , V CRF , ECRF ) dY ea − Q(Y ea ) log Q(Yea )dY
ea ,
̸= a

where ̸= m means all indices except the mth . Now, we take the derivative of the ELBO with respect to
Q (Y
ea ):

d ELBO h i
= EQ̸=a log P(Y
ea |Y̸=a , V CRF , ECRF ) − log Q(Y
ea ) − 1 = 0.
dQ(Yea )
Therefore, n h  io
ea ) ∝ exp E−a log P Y
∀ a ∈ V CRF , Q(Y e|Y, V CRF , ECRF .

D.2 A S Y M P T O T I C B E H AV I O R O F C R F N E I G H B O R H O O D S I Z E

D.2.1 Proof of Lemma 6.3.3

n ( n +1)
Lemma 6.3.3 For any integer r in {0, . . . , 2 }, the number of CRF neighboors N CRF ( a) for any
a ∈ V CRF , i.e., the set of graphs [A,
e Xe ] with a Hamming distance smaller or equal than r, for each a ∈ V CRF ,
we have the following lower bound:

2 H (ϵ)n(n+1)/2
p ≤ N CRF ( a) , (d.2)
4n(n + 1)ϵ(1 − ϵ)
2r
where 0 ≤ ϵ = n ( n +1)
≤ 1 and H (·) is the binary entropy function, i.e., H (ϵ) = −ϵ log2 (ϵ) − (1 −
ϵ) log2 (1 − ϵ).

Proof. We use Stirling’s formula;


√ s −s

1 1

∀s, s! = 2πss e exp − +... . (d.3)
12s 360s3
n ( n +1) n ( n +1)
Thus, for r in {0, . . . , 2 } and L = 2 , we write
168 A P P E N D I X : A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

 n ( n +1)   
2 L
= (d.4)
r r
L
= (d.5)
r!( L − r )!

2πLL L e− L exp [−1/12r − 1/12( L − r )]
≥√ p . (d.6)
2πrrr e−r 2π ( L − r )( L − r )( L−r) e−( L−r)

For L ≥ 4, for any r ∈ {1, . . . , L}, we always have L − r ≥ 3 or r ≥, thus,


1 1 1 1 1
+ ≤ + = . (d.7)
12r 12( L − r ) 12 36 9
Therefore,
1√
 
1 1 1 1
exp − − ≥ + = e−1/9 ≥ π. (d.8)
12r 12( L − r ) 12 36 2
n ( n +1)
We can then therefore derive a lower bound for ( r
2 ):
 n ( n +1)   
2 L
= (d.9)
r r
√ √
2πLL L e− L 21 π
≥√ p (d.10)
2πrrr e−r 2π ( L − r )( L − r )( L − r )e−( L−r)

LL L
=√ p (d.11)
8rrr ( L − r )( L − r )( L−r)
s
L LL
= (d.12)
8r ( L − r ) rr ( L − r )( L−r)
1 1
=p r (1 − ϵ ) L −r
(d.13)
8Lϵ(1 − ϵ) ϵ
1
=p ϵ− Lϵ (1 − ϵ) L(1−ϵ) (d.14)
8Lϵ(1 − ϵ)
1
=p 2 LH (ϵ) . (d.15)
8Lϵ(1 − ϵ)
n ( n +1)
For d in {0, . . . , 2 }, the number of possible graphs with a Hamming distance equal to d from a graph
a = [ G, x ] is ( Ld). Thus,
r  
L
N CRF
( a) = ∑ (d.16)
d =0
d
 
L
≥ (d.17)
r
1
≥p 2 LH (ϵ) (d.18)
8Lϵ(1 − ϵ)
1
=p 2 H (ϵ) N ( N +1)/2 . (d.19)
4N ( N + 1)ϵ(1 − ϵ)
D.3 T H E S E T O F H Y P E R PA R A M E T E R S U S E D T O C O N S T R U C T T H E C R F 169

D.2.2 Empirical Investigation of the Lower Bound

We empirically investigate the evolution of the lower-bound stated in Lemma 6.3.3 as a function of ratio
ϵ = ϵ(r ). As noticed, the number of neighbors increases exponentially as the radius increases. This
motivates the need for sampling strategies to reduce the size of the CRF.
The evolution of the lowerbound when the number of nodes is n = 500
10140
10123
10106
Lower Bound

1089
1072
1055
1038
1021

0.0 0.1 0.2 0.3 0.4 0.5

Figure d.1 – The effect of the radius on the lower bound stated in Lemma 6.3.3.

D.3 T H E S E T O F H Y P E R PA R A M E T E R S U S E D T O C O N S T R U C T T H E C R F

In Tables d.1 and d.2, we present the hyperparameters used to run the inference of RobustCRF for
each dataset. The number of iterations was fixed to 2. pr and we compute the radius as a floor function,
i. e. r = ⌊ pr × m⌋, where m = | E| is the number of existing edges in the original graph.

Table d.1 – The optimal RobustCRF’s hyperparameters for the feature-based attacks.

Hyperparameter Cora CiteSeer PubMed CS

r 0.1 0.9 0.3 0.3


σ 0.9 0.8 0.9 0.5

Table d.2 – The optimal RobustCRF’s hyperparameters for the structure-based attacks.

Hyperparameters Cora CoraML CiteSeer PolBlogs

pr 0.02 0.02 0.04 0.005


σ 0.05 0.2 0.05 0.05
170 A P P E N D I X : A P O S T- H O C A P P R O A C H W I T H C O N D I T I O N A L R A N D O M F I E L D S

Table d.3 – Statistics of the node classification datasets used in our experiments.
Dataset #Features #Nodes #Edges #Classes

Cora 1,433 2,708 5,208 7


CoraML 300 2,995 8,226 7
CiteSeer 3,703 3,327 4,552 6
PubMed 500 19,717 44,338 3
CS 6,805 18,333 81,894 15
PolBlogs - 1,490 19,025 2
Texas 1,703 183 309 5
OGBN-Arxiv 128 31,971 71,669 40

D.4 E X P E R I M E N TA L S E T U P

Datasets. For our experiments, we focus on node classification within the general perspective of node
representation learning. We use the citation networks Cora, CoraML, CiteSeer, and PubMed (Sen et al.,
2008). We additionally consider the co-authorship network CS (Shchur et al., 2018b), the blog and citation
graph PolBlogs (Adamic and Glance, 2005), and the non-homophilous dataset Texas (Lim et al., 2021b).
More details and statistics of the datasets can be found in Table d.3. For the CS dataset, we randomly
selected 20 nodes from each class to form the training set and 500/1000 nodes for the validation and
test sets (Yang, Cohen, and Salakhudinov, 2016). For all the remaining datasets, we adhere to the public
train/valid/test splits provided by the datasets.
Implementation Details. We used the PyTorch Geometric (PyG) open-source library, licensed under
MIT (Fey and Lenssen, 2019). Additionally, for adversarial attacks in this study, we used the DeepRobust
package 1 . The experiments were conducted on an RTX A6000 GPU. For the structure-based CRF, we
leveraged the sampling strategy detailed in Section 6.3. The set of hyperparameters for each dataset can
be found in Appendix d.3. We compute the similarity gab between two inputs a = [ G, X ] and b = [ G̃, X̃ ]
using the cosine similarity for the features based attacks, namely CosSim( X, X̃ ), while for the structural
attacks, we use the prior distribution gab = (dr ) 21r , where d is the value of the Hamming distance between
the original graph a and its sampled neighbor b. We note that our code is provided in the supplementary
materials and will be made public upon publication.

D.5 TIME AND COMPLEXITY OF ROBUSTCRF

In Table d.4, we present the average time (in seconds) required for RobustCRF inference. As observed,
the inference time increases exponentially with larger values of K and L. Nevertheless, empirical evidence
suggests that a small number of samples is sufficient to improve GNN robustness. In our experiments, we
specifically used L = 5 samples and set the number of iterations to 2.

D.6 RESULS ON OGBN-ARXIV

To further evaluate the effectiveness of RobustCRF, we conducted experiments on the OGBN-Arxiv


dataset, a large-scale Open Graph Benchmark (OGB) dataset commonly used for evaluating node classifi-
cation in graph-based learning models. The dataset consists of a citation network where nodes represent

1. [Link]
D.6 R E S U L S O N O G B N - A R X I V 171

Table d.4 – Inference time for different values of the number of iterations/samples.

Num Samples L 0 Iter 1 Iter 2 Iter

5 0.26 ± 0.52 1.99 ± 0.54 16.40 ± 2.04


10 0.20 ± 0.41 3.24 ± 0.47 58.28 ± 1.15
20 0.22 ± 0.44 5.86 ± 0.53 224.86 ± 0.71

Table d.5 – Attacked classification accuracy on the OGBN-Arxiv dataset.

Dataset GCN NoisyGCN RobustCRF

Clean 60.41±0.15 59.97±0.11 60.28±0.15


Random 58.97±0.24 58.71±0.10 59.03±0.26
PGD 50.24±0.42 50.26±0.37 50.10±0.56

ArXiv papers and edges denote citation links Hu et al., 2020c. In Table d.5, we present the attacked
classification accuracy of the GCN, a baseline and the proposed RobustCRF on the OGBN-Arxiv dataset.
A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA
e
A U G M E N TAT I O N

173
174 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

E .1 PROOF OF THEOREM 7.3.1

In this section, we provide a detailed proof of Theorem 7.3.1, aiming to derive a theoretical upper bound
for both the generalization gap and the Rademacher complexity.
Theorem 7.3.1 Let ℓ be a classification loss function with LLip as a Lipschitz constant and ℓ(·, ·) ∈ [0, 1].
Then, with a probability at least 1 − δ over the samples Dtrain , we have,
r
2 log(4/δ) h i
EG∼D ℓ(G , θ̂aug ) − EG∼D [ℓ(G , θ⋆ )] ≤ 2R(ℓaug ) + 5 + 2LLip EG∼D ,G∼
 
e Aλ G−G .
e
N
Moreover, we have, h i
R(ℓaug ) ≤ R(ℓ) + max LLip EGem ∼ Aλ Genm − Gn .
n n

Proof. We will decompose EG∼D ℓ(G , θ̂aug ) − EG∼D [ℓ(G , θ⋆ )] into a finite sum of 5 terms as follows,
 

EG∼D ℓ(G , θ̂aug ) − EG∼D [ℓ(G , θ⋆ )] = u1 + u2 + u3 + u4 + u5


 

where,
h h ii
u1 = EG∼D ℓ(G , θ̂aug ) − EG∼D EG∼
 
e Aλ ℓ( G , θ̂aug ) ,
e
h h ii 1 N h i
N n∑
u2 = EG∼D EG∼ e Aλ ℓ( G
e , θ̂ aug ) − E Genm ∼ Aλ ℓ( G
e m
n , θ̂ aug ) ,
=1
1 N h i 1 N h i
N n∑ N n∑
u3 = E m
Gn ∼ Aλ
e ℓ( G
e m
n , θ̂ aug ) − E m
Gn ∼ Aλ
e ℓ( G
e m
n , θ̂ ⋆ ) ,
=1 =1
1 N h i h h ii
u4 = ∑ EGem ∼ Aλ ℓ(Genm , θ̂⋆ ) − EG∼D EG∼
N n =1 n
e Aλ ℓ( G
e, θ⋆ ) ,
h h ii
u5 = EG∼D EG∼ e Aλ ℓ( G , θ⋆ )
e − EG∼D [ℓ(G , θ⋆ )] .

We upperbound each of the terms in the sum. We get,


h h ii h h ii
u1 + u5 = EG∼D ℓ(G , θ̂aug ) − EG∼D EG∼ + EG∼D EG∼ − EG∼D [ℓ(G , θ⋆ )]
 
e Aλ ℓ( G , θ̂aug ) e Aλ ℓ( G , θ⋆ )
e e
h h ii h h ii
≤ EG∼D ℓ(G , θ̂aug ) − EG∼D EG∼ + EG∼D EG∼ − EG∼D [ℓ(G , θ⋆ )]
 
e Aλ ℓ( G , θ̂aug ) e Aλ ℓ( G , θ⋆ )
e e
h h ii
≤ 2 sup EG∼D EG∼ e Aλ ℓ( G , θ )
e − EG∼D [ℓ(G , θ )]
θ ∈Θ
h h i i
≤ 2 sup EG∼D EG∼ e Aλ ℓ( G
e , θ ) − ℓ(G , θ )
θ ∈Θ
h h ii
≤ 2 sup EG∼D EG∼ e Aλ ℓ( G , θ ) − ℓ(G , θ )
e
θ ∈Θ
h i
≤ 2LLip sup EG∼D EG∼
e Aλ Ge − G .
θ ∈Θ

For the term u4 , we apply McDiarmid’s inequality. Since the classification loss satisfy ℓ(·) ∈ [0, 1], we get
for k ∈ {0, . . . , N },
For all {(Gn , yn )}nN=1 , {(G ′ n , y′n )}nN=1 , θ, such that ∀n ̸= k, Gn = G ′ n and Gk ̸= G ′ k :
E .1 P R O O F O F T H E O R E M 7.3.1 175

N N N
1 1 1
∑ EG∼ ∑ EG∼ ′
∑ EG∼ ℓ(Gn , θ ) − ℓ(Gn′ , θ )
   
e Aλ [ℓ(Gn , θ )] − e Aλ ℓ(Gn , θ ) = e A
N n =1
N n =1
N n =1
λ

1 ′
EG∼
 
≤ e Aλ ℓ(Gk , θ ) − ℓ(Gk , θ )
N
≤ 2/N.

The first equality is obtained by your claim that ∀n ̸= k, Gn = G ′ n and Gk ̸= G ′ k , the last inequality is
obtained by the fact that ℓ(·) ∈ [0, 1].
Thus,
!
1 N h i h h ii
N n∑
∀t > 0, P (u4 ≥ t) = P E Genm ∼ Aλ ℓ( G
e m
n , θ ⋆ ) − E G∼D E G∼
e Aλ ℓ( G
e , θ ⋆ ) ≥t
=1
2t2
 
≤ exp − N
∑n=1 4/N 2
Nt2
 
= exp − .
2
 2

Therefore, for δ ∈]0, 1], and for t = 2 log(1/δ)/N, i.e. exp − Nt2 = δ., we have,
p

 q 
P u4 ≥ 2 log(1/δ)/N ≤ δ.

Therefore,
r ! r !
2 log(1/δ) 2 log(1/δ)
P u4 < = 1 − P u4 ≥ ≥ 1 − δ.
N N
Thus, with a probability of at least 1 − δ,
r r
2 log(1/δ) 2 log(4/δ)
u4 ≤ < .
N N
Moreover, Rademacher complexity holds for u2 ,

r
N
h h ii 1 h i 2 log(4/δ)
u2 = EG∼D EG∼
e Aλ ℓ( G , θ̂aug )
e −
N ∑ EGem
n ∼ Aλ
ℓ(Genm , θ̂aug ) ≤ 2R(ℓaug ) + 4
N
.
n =1
h h ii
The above inequality tells us that the true risk EG∼D EG∼ e Aλ ℓ( G , θ̂aug )
e is bounded by the empirical
h i
risk N1 ∑nN=1 EGem ∼ Aλ ℓ(Genm , θ̂aug ) plus a term depending on the Rademacher complexity of the augmented
n
hypothesis class and an additional term that decreases with the size of the sample h N. i
Additionally, since θ̂aug is the optimal parameter for the loss N ∑n=1 EGem ∼ Aλ ℓ(Gn , θ̂ ) , thus,
1 N em
n

u3 ≤ 0.

By summing all the inequalities, we conclude that,


r
2 log(4/δ) h i
EG∼D ℓ(G , θ̂aug ) − EG∼D [ℓ(G , θ⋆ )] < 2R(ℓaug ) + 5 + 2LLip EG∼D EGem ∼ Aλ Genm − Gn .
 
N n
176 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

Part 2 of the proof.


" #
N N
1 1
R(ℓaug ) − R(ℓ) = Eϵn ∼ Pϵ sup
N ∑ ϵn ℓaug (Gn , θ ) − sup N ∑ ϵn ℓ(Gn , θ )
θ ∈Θ n =1 θ ∈Θ n =1
" #
N N
1 1
≤ Eϵn ∼ Pϵ sup ∑ ϵn ℓaug (Gn , θ ) − N ∑ ϵn ℓ(Gn , θ )
θ ∈Θ N n =1 n =1
" #
N
1
= Eϵn ∼ Pϵ sup ∑ ϵn (ℓaug (Gn , θ ) − ℓ(Gn , θ ))
θ ∈Θ N n =1
" #
N
1
≤ Eϵn ∼ Pϵ sup ∑ |ϵn (ℓaug (Gn , θ ) − ℓ(Gn , θ ))|
θ ∈Θ N n =1
N
1
≤ sup
N ∑ |ℓaug (Gn , θ ) − ℓ(Gn , θ )|
θ ∈Θ n =1
N
1 h i
= sup
N ∑ EGnm ∼ Aλ ℓ(Gnλ , θ ) − ℓ(Gn , θ )
θ ∈Θ n =1
h i
≤ max LLip EGnm ∼ Aλ Genm − Gn .
n∈{1,...,N }

E .2 PROOF OF PROPOSITION 7.3.2

Proposition 7.3.2 Let δD denote the discrete distribution of the training graph representations. Suppose
we sample new augmented graph representations from a distribution Qλ defined on Rd . Then, the following
inequality holds,
h i √ q √ 
Eh∼δ ,he ∼Q ∥h − h∥ ≤ 2 · sup ∥h − h∥
e e KL(δD ∥ Qλ ) + 2 .
D λ
h∼δD
e ∼ Qλ
h

where KL(δD ∥ Qλ ) is is the Kullback-Leibler divergence from δD to Qλ .

Proof. We have,

h i Z Z
Eh∼δG ,he ∼Q ∥h − h∥ =
e ∥h − h
e ∥δD (h) Qλ (h
e ) dh dh,
e
λ
h ∈ Rd h ∈ Rd
e

where δD : h 7→ 1
N ∑nN=1 δhGn (h), and δ is the Dirac distribution.

h i Z Z
Eh ∼ δ ∥h − h∥ =
e ∥h − h
e ∥δD (h) Qλ (h
e ) dh dh
e
D ,h∼ Qλ h ∈ Rd h ∈ Rd
e
e
Z Z
≤C δD (h) − Qλ (h
e ) δD (h) Qλ (h
e ) dh dh,
e
Rd Rd

where
∥h − h
e ∥δD (h) Qλ (h
e)
C = sup .
h∼δD δD (h) − Qλ (he)
h∼Q
e
h̸=e
h
E .2 P R O O F O F P R O P O S I T I O N 7.3.2 177

h i Z Z
Eh ∼ δ ∥h − h∥ ≤ C
e δD (h) − Qλ (h
e ) dh dh
e
D , h ∼Q h ∈ Rd h ∈ Rd
e
e
Z Z Z Z
≤C δD (h) − Qλ (h
e ) dh dh
e +C
h ∈ Rd
δD (h) − Qλ (h
e ) dh dh
e.
h ∈ Rd h ∈ Rd
e
h ∈ Rd
e
h=e h h̸=eh

We first get an upperbound for the first term in the right side of the inquality,
Z Z Z Z 
δD (h) − Qλ (h) dh dh =
e e δD (h) − Qλ (h) dh mathopdh
e e
h ∈ Rd eh∈Rd h=e
h h ∈ Rd h∈Rd h=e
e h
Z Z 
= |δD (h) − Q(h)| dh dh
e
h ∈ Rd h∈Rd h=e
e h
Z  Z 
= δh (h) dh |δD (h) − Q(h)| dh
e e
h ∈ Rd h ∈ Rd
e
Z
= |δD (h) − Q(h)| dh
h ∈ Rd
rZ
≤ |δD (h) − Q(h)|2 dh, using Jensen’s Inequaliy.
h ∈ Rd
R
The term h∈Rd |δD (h) − Q(h)| dh corresponds to the Total Variation Distance (TD) d TV between PG and
Q (Yao and Liu, 2024), i.e.
Z
d TV ( PG , Q) = 2 sup | PG ( A) − Q( A)| = | PG ( x ) − Q( x )| dx = ∥ PG − Q∥1 .
A ⊂Rd R

Using Pinsker’s Inequality, we have,


Z q
|δD (h) − Q(h)| dh ≤ 2KL(δD ∥ Qλ )
h ∈ Rd

For the second term in the upperboud, we have,


!
N N
1 1
Z Z Z

h ∈ Rd h ∈ Rd
e δD (h) − Qλ (h
e ) δD (h) dh dh
e=
h ∈ Rd
e
N ∑ δhGn (h) − Qλ (h
e)
N ∑ δhGn (h) dh dh
e
h̸=eh h̸=eh n= N n= N
N
1 1
Z
=
N ∑ h ∈ Rd
e
N
− Qλ (h e ≤ 2,
e ) dh dh
n =1 e
h ̸ = hGn

because distributions are bounded between 0 and 1


Let now upperbound the constant C. We have ∀ a, b ∈ R+ , ab ≤ 21 ( a − b)2 . Therefore,
178 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

∥h − h
e ∥δD (h) Qλ (h
e)
C = sup
h∼δD δD (h) − Qλ (he)
h∼Q
e
h̸=e
h
1
≤ sup ∥h − h
e ∥ δD (h) − Qλ (h
e)
2 h∼δD
h∼Q
e
h̸=e
h
1
≤ sup ∥h − h
e ∥2 Because distributions are bounded between 0 and 1
2 h∼δD
h∼Q
e
h̸=e
h

= sup ∥h − h
e∥
h∼δD
h∼Q
e

Therefore,
h i √ q √ 
Eh∼δD Eeh∼Q ∥h − h∥ ≤ 2 sup ∥h − h∥
e e KL(δD ∥ Qλ ) + 2 .
h∼δD
h∼Q
e

E .3 PROOF OF THEOREM 7.3.4

In this section, we present the detailed proof of Theorem 7.3.4, which allows us to e perform an in-depth
theoretical analysis of our augmentation strategy through the lens of influence functions.
Theorem 7.3.4 Given a test graph Gk from the test set, let θ̂ = arg minθ L be the GNN parameters
that minimize the objective function in (7.1). The impact of upweighting the objective function L to
Laug em em
n,m = L + ϵn,m ℓ( Gn , θ ), where Gn is an augmented graph candidate of the training graph Gn and ϵn,m is
a sufficiently small perturbation parameter, on the model performance on the test graph Gktest is given by

dℓ(Gktest , θ̂ϵn,m )
= −∇θ ℓ(Gktest , θ̂ )H− 1
∇θ ℓ(Genm , θ̂ ),
dϵn,m θ̂

aug aug
where θ̂ϵn,m = arg minθ Ln,m denotes the parameters that minimize the upweighted objective function Ln,m
and Hθ̂ = ∇2θ L(θ̂ ) is the Hessian Matrix of the loss w.r.t. the model parameters.

Proof. Let Genm be an augmented graph candidate of the training graph Gn and ϵn,m is a sufficiently small
perturbation parameter. The parameters θ̂ and θ̂ and θ̂ϵn,m the parameters that minimize the empirical risk
on the train set, i.e.,

θ̂ = arg min L,
θ
aug
θ̂ϵn,m = arg min Ln,m = arg min L + ϵn,m ℓ(Genm , θ ).
θ θ

Therefore, we examine its first-order optimality conditions,


E .4 M AT H E M AT I C A L E X P R E S S I O N S O F G C N A N D G I N 179

0 = ∇θ̂ L (e.1)
 
0 = ∇θ̂ϵ
n,m
L + ϵn,m ℓ(Genm Gnm , θ ) . (e.2)

Using Taylor Expansion, we now develop Eq. (e.2). We have limϵn,m →0 θ̂ϵn,m = θ̂, thus,
h i h i
m 2 2 m

0 ≃ ∇θ̂ L(θ̂ ) + ϵn,m ∇θ̂ ℓ(Gn , θ̂ ) + ∇θ̂ L(θ̂ ) + ϵn,m ∇θ̂ ℓ(Gn , θ̂ ) θ̂ϵn,m − θ̂ .
e e

Therefore,
h i −1 h i
θ̂ϵn,m − θ̂ = − ∇2θ̂ L(θ̂ ) + ϵn,m ∇2θ̂ ℓ(Genm , θ̂ ) ∇θ̂ L(θ̂ ) + ϵn,m ∇θ̂ ℓ(Genm , θ̂ ) .

Dropping the ◦(ϵn,m ) terms, and using the Equation e.1, i.e. ∇θ̂ L = 0, we conclude that,

θ̂ϵn,m − θ̂ h i −1
= − ∇2θ̂ L(θ̂ ) ∇θ̂ ℓ(Genm , θ̂ ).
ϵn,m

Therefore,

dθ̂ϵn,m θ̂ϵ − θ̂ h i −1
≃ n,m = − ∇2θ̂ L(θ̂ ) ∇θ̂ ℓ(Genm , θ̂ ).
dϵn,m ϵn,m

dℓ(Gktest , θ̂ϵn,m ) dℓ(Gktest , θ̂ϵn,m ) dθ̂ϵn,m


=
dϵn,m dθ̂ϵn,m dϵn,m
= −∇θ ℓ(Gktest , θ̂ )H−
θ̂
1
∇θ ℓ(Genm , θ̂ ).

E .4 M AT H E M AT I C A L E X P R E S S I O N S O F G C N A N D G I N

In this section, we provide concise definitions of two widely used GNN architectures: Graph Convolutional
Networks (GCN) and Graph Isomorphism Networks (GIN). These architectures differ in how they aggregate
and combine information from neighboring nodes in a graph.

( G C N ) K I P F A N D W E L L I N G , 2 0 1 7 C . The GCN updates


G R A P H C O N VO L U T I O N A L N E T WO R K
node embeddings by aggregating normalized features from their neighbors. Specifically, for a node v ∈ V ,
(t)
its feature vector hv at layer t is computed as,
!
1
∑ pdeg(v) deg(u) W(t) hu
(t) ( t − 1 )
hv = σ ,
u∈N (v)∪{v}

( t −1)
where hu is the feature vector of node u at layer t − 1, W(t) is a trainable weight matrix for layer t, and
σ (·) is a non-linear activation function, such as ReLU.
180 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

GRAPH ISOMORPHISM NETWORK ( G I N ) X U E T A L . , 2 0 1 9 B . GIN is designed to match the ex-


pressiveness of the Weisfeiler-Lehman (WL) graph isomorphism test. It updates a node’s embedding by
aggregating its own feature with those of its neighbors, followed by a a multi-layer perceptron (MLP). The
update rule at each message passing layer t is
 
+ ∑ hu
(t) ( t −1) ( t −1)
hv = MLP(t) (1 + ϵ(t) ) · hv ,
u∈N (v)

where ϵ(t) is a fixed or learnable scalar. GIN allows for more expressive feature transformations compared
to methods with fixed aggregation schemes, enabling it to distinguish a broader range of graph structures.

(t)
M AT R I X F O R M S A N D C O M PA R I S O N .
The node representations hv at each layer can be expressed in
matrix form as H(t) ∈ R p×dt , where p isthe number of nodes in the graph, d is hidden dimension in the
(t)
t−th layer, and H(t) is the concatenation of all node representations hv for v ∈ V . For GCN, the update
rule can be written as:
 
H(t) = σ De −1/2 A
eDe −1/2 H(t−1) W(t) ,

where Ae = A + I is the adjacency matrix with self-loops, and D


e is its diagonal degree matrix. In contrast,
for GIN, the update rule is given by:
 
H(t) = MLP(t) (1 + ϵ(t) )H(t−1) + AH(t−1) .

If the same MLP is used, comprising a learnable linear layer followed by a ReLU activation, the only
difference between GCN and GIN lies in the graph shift operator : GCN uses the degree-normalized
operator D e −1/2 A
eDe −1/2 , while GIN uses (1 + ϵ(t) )I + A.
Therefore, when node features are taken as constant and only the graph structure is considered, the
difference in their Lipschitz behavior can be traced back to the function that maps the adjacency matrix A
(and degrees) to these respective shift operators. In other words, the Lipschitz constant difference arises
from whether the adjacency information is normalized (D e −1/2 A
eDe −1/2 ) or as ((1 + ϵ(t) )I + A).

L I P S C H I T Z C O N S TA N T S O F G C N A N D G I N .
Under constant node features, the Lipschitz constant of
each architecture can be decomposed into two parts: one accounting for the sensitivity of the mapping
from the adjacency matrix to the respective graph shift operator, i.e., degree-normalized De −1/2 A
eDe −1/2 for
GCN vs. (1 + ϵ)I + A for GIN, and the other capturing all shared learnable transformations.
Let ℓGCN and ℓGIN be respectively the graph shift operators A 7→ D e −1/2 A
eDe −1/2 and A 7→ (1 + ϵ)I + A,
and let ℓparams represent the product of Lipschitz factors arising from the shared functions. Then,

LGCN = ℓGCN × ℓparams


LGIN = ℓGIN × ℓparams

It becomes clear that the difference between GCN and GIN Lipschitz constants depends solely on
whether the graph adjacency is normalized, since ℓparams is common to both. It’s straightforward that
ℓGIN ≤ 1 since the corresponding graph shift operator is just a translation. Let now derive an upperbound
for ℓGIN . Let A1 , A2 ∈ R p× p be adjacency matrices of graphs on n nodes. For each adjacency matrix Ai ,
E .5 C O N F I G U R AT I O N M O D E L S 181

(i ) (i )  (i )
we define the diagoanl digree matrix, Di = diag deg1 , . . . , deg p where ∀ j ≤ p, deg j = ∑nk=1 (Ai ) jk .
We have,
1 1 1 1 1 1 1 1 1 1 1 1
−2 −2 −2 −2 −2 −2 − − − − − −
D1 A1 D1 − D2 A2 D2 = D1 A1 D1 − D1 2 A2 D1 2 + D1 2 A2 D1 2 − D2 2 A2 D2 2
1  −1  −1 1 1 1  −1 1
−2 −  − − − 
= D1 A1 − A2 D1 2 + D1 2 − D2 2 A2 D1 2 + D2 2 A2 D1 2 − D2 2 .
| {z } | {z }
(i) (ii)

We can upperbound each of (i), and (ii). For (i), we have the following upperbound,
1  −1 1
−2 − 1
D1 A1 − A2 D1 2 ≤ ∥ D1 2 ∥2 ∥ A1 − A2 ∥ ≤ ∥ A1 − A2 ∥ ,
δ1,min
(i )
where δi,min = min j deg j . For the second (ii), if we consider for example ∥∥ as the L1 norm, then,

∥(ii)∥ ≤ ∥D1−1/2 − D2−1/2 ∥∥A2 ∥ ∥D1−1/2 ∥ + ∥D2−1/2 ∥ ,




2
≤ ∥D1−1/2 − D2−1/2 ∥∥A2 ∥
min(δ1,min , δ2,min )1/2
2M
≤ 1/2
∥D1−1/2 − D2−1/2 ∥
min(δ1,min , δ2,min )
2M
≤ ∥ D1 − D2 ∥
min(δ1,min , δ2,min )5/2
2M
≤ ∥ A1 1 p − A2 1 p ∥
min(δ1,min , δ2,min )5/2
2Mp
≤ ∥ A1 − A2 ∥ ,
min(δ1,min , δ2,min )5/2

where M is the maximum norm of the adjacency matrix in the graph dataset, and 1 p ∈ R p is the vector
of ones.
Putting both inequalities together, we can come up with an upperbound for ℓGCN that depends on the
minimum degree in the dataset.

E .5 C O N F I G U R AT I O N M O D E L S

In this section, we present a novel adaptation of configuration models as a graph data augmentation
technique for GNN. Configuration models Newman, 2013 enable the generation of randomized graphs
that maintain the original degree distribution. We can, therefore, leverage this strategy to improve the
generalization of GNNs. Below, we present the steps involved in our approach to using Configuration
Models for Graph Data Augmentation:
1. Extract edges: For each training graph Gn , we first extract the complete set of edges En .
2. Stub creation: Using a Bernoulli distribution with parameter r ∈ [0, 1], we randomly select a subset
of candidate edges and break them to create stubs (half-edges).
3. Stub pairing: We then randomly pair these stubs to form new edges, creating a randomized graph
structure with the same degree distribution.
182 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

Table e.1 – Classification accuracy (± standard deviation) on different benchmark node classification datasets for the
data augmentation baselines based on the GIN backbone. The higher the accuracy (in %) the better the
model.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

Config Models w/ GCN 71.70±3.16 48.40±3.88 74.97±6.77 70.08±4.93 69.01±3.44


Config Models w/ GIN 71.70±4.24 49.00±3.44 81.43±10.05 68.34±5.30 71.61±5.93

Table e.1 shows the performance of this approach on the two GNN backbones, GCN and GIN.
As noticed, the configuration model-based graph augmentation method performs competitively with the
baselines and even outperforms them in certain cases. This underscores the importance of Theorem
7.3.1. When compared to our approach GRATIN, the latter gives better results across different datasets
and GNN backbones. This difference is primarily due to the configuration model based approach being
model-agnostic, whereas GRATIN leverages the model’s weights and architecture, as explained in Section
7.3.4 and supported by Theorem 7.3.4.

E .6 A B L AT I O N S T U DY

To provide additional comparison and motivate the use of GMMs with the EM algorithm within GRATIN,
we expanded our evaluation to include additional methods for modeling the distribution of the graph
representations. Specifically, the comparison includes:
— GMM w/ Variational Bayesian Inference (VBI): We specifically compared the Expectation-Maximization
(EM) algorithm, discussed in the main paper, with the Variational Bayesian (VB) estimation technique
for parameter estimation of each Gaussian Mixture Model (GMM) (Tzikas, Likas, and Galatsanos,
2008) for both the GCN and GIN models. The objective of including this baseline is to explore
alternative approaches for fitting GMMs to the graph representations.
— Kernel Density Estimation (KDE): KDE is a Neighbor-Based Method and a non-parametric approach
to estimating the probability density (Härdle et al., 2004). KDE estimates the probability density
function by placing a kernel function (e.g., Gaussian) at each data point. The sum of these kernels
approximates the underlying distribution. Sampling can be done using techniques like Metropolis-
Hastings. The purpose of using KDE as a baseline is to evaluate alternative distributions different
from the Gaussian Mixture Model (GMM).
— Copula-Based Methods: We model the dependence structure between variables using copulas,
while marginal distributions are modeled separately. We sample from marginal distributions and then
transform them using the copula (Nelsen, 2006).
— Generative Adversarial Network (GAN): GANs are powerful generative models that learn to
approximate the data distribution through an adversarial process between two neural networks.
To evaluate the performance of deep learning-based generative approaches for modeling graph
representations, we included tGAN, a GAN architecture specifically designed for tabular data (Yang,
Lai, and Lin, 2012). We particularly train tGAN on the graph representations and then sample new
graph representations from the generator.
We compare these approaches for both the GCN and GIN models in Tables e.2 and e.3, respectively. As
noticed, GMM with EM consistently outperforms the alternative methods across most datasets in terms of
accuracy. The VBI method, an alternative approach for estimating GMM parameters, yields comparable
E .7 T R A I N I N G A N D A U G M E N TAT I O N T I M E 183

Table e.2 – Ablation study on the density estimation scheme for learned GCN representations in GRATIN.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

GMM w/ EM 71.00±4.40 49.82±4.26 76.05±6.47 70.97±5.07 71.90±2.81


GMM w/ VBI 71.00±4.21 49.53±4.26 76.05±6.47 70.97±4.52 71.64±2.90
KDE 55.90±10.29 39.53±2.87 66.64±6.79 59.56±2.62 58.66±3.97
Copula 69.80±4.04 47.13±3.45 74.44±6.26 65.04±3.37 65.70±3.04
GAN 70.60±3.41 48.80±5.51 75.52±4.96 69.98±5.46 66.26±3.72

Table e.3 – Ablation study on the density estimation scheme for learned GIN representations in GRATIN.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

GMM w/ EM 71.70±4.24 49.20±2.06 88.83±5.02 71.33±5.04 68.61±4.62


GMM w/ VBI 71.40±2.65 47.80±2.22 88.30±5.19 70.25±4.65 67.82±4.96
KDE 69.10±3.93 41.46±3.02 77.60±6.83 60.37±3.04 67.48±6.18
Copula 70.60±2.61 47.60±2.29 88.30±5.19 70.16±4.55 67.91±4.90
GAN 70.50±3.80 48.40±1.71 88.83±5.02 71.33±5.55 67.74±4.82

performance to the EM algorithm. This consistency across datasets highlights the effectiveness and
robustness of GMMs in capturing the underlying data distribution.
In certain cases, particularly with the GIN model, we observed competitive performance from the GAN
approach, which, unlike GMM, requires additional training. Hence, GMMs provide a more straightforward
and efficient solution.

E .7 T R A I N I N G A N D A U G M E N TAT I O N T I M E

We compare the data augmentation times of our approach and the baselines in Table e.4. In addition
to outperforming the baselines on most datasets, our approach offers an advantage in terms of time
complexity. The training time of baseline models varies depending on the augmentation strategy used,
specifically, whether it involves pairs or individual graphs. Even in cases where a graph augmentation has
a low computational cost for some baselines, training can still be time-consuming as multiple augmented
graphs are required to achieve satisfactory test accuracy. For instance, methods like DropEdge, DropNode,
and SubMix, while computationally simple, require generating multiple augmented samples at each epoch,
thereby increasing the overall training time. Following the framework of Yoo, Shim, and Kang, 2022, we
must sample several augmented graphs for each training graph at every epoch to achieve optimal results.
In contrast, GRATIN introduces a more efficient approach by generating only one augmented graph per
training instance, which is reused across all epochs. This design ensures a balance between computational
efficiency and augmentation effectiveness, reducing the overall training burden while maintaining strong
performance. The only baseline that is more time-efficient than our approach is GeoMix; however, our
method consistently outperforms GeoMix across all settings, as shown in Tables 7.1 and 7.2.
Furthermore, the performance of simple augmentation baselines such as DropEdge, DropNode, and
SubMix significantly drops when using only one augmentation per training graph for all the epochs, which
is the framework adopted in GRATIN, as shown in Tables e.5 and e.6. Tables e.5 and e.6, these baselines
184 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

Table e.4 – Mean training and augmentation time in seconds of our model in comparison to the other benchmarks.

Time Model IMDB-BIN MUTAG DD

Vanilla - - -
DropEdge 0.02 0.01 0.01

1 DropNode 0.01 0.02 0.01
Aug. Time SubMix 1.27 0.23 0.45
G -Mixup 0.74 0.11 4.26
GeoMix 2,344.12 73.52 1,005.35
GRATIN 2.87 0.51 3.25
Vanilla 765.96 99.32 428.10
DropEdge 892.14 596.82 3,037.30

2 DropNode 884.71 803.63 3,325
Train. Time SubMix 1,711.01 1,487.03 2,751.92
G -Mixup 148.71 28.14 177.55
GeoMix 89.01 101.82 123.41
GRATIN 774.47 101.56 438.39

Table e.5 – Classification accuracy (± std) on different benchmark datasets for the data augmentation baselines
using the GCN backbone. Higher accuracy (in %) is better.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

DropEdge 71.40±3.69 48.47±2.36 73.88±7.98 67.56±4.50 66.04±4.35


DropNode 71.80±4.11 48.60±3.45 72.78±8.01 67.83±4.75 66.55±3.97
SubMix 72.50±3.90 47.93±3.58 71.72±10.59 59.74±2.83 62.25±3.29

experience a significant drop in accuracy under this constraint, emphasizing their reliance on generating
diverse augmentations at each epoch to maintain strong performance, c.f. the results in Tables 7.1 and 7.2.
The only graph augmentation baseline with comparable or better time complexity than GRATINis G -Mixup.
However, GRATINconsistently outperforms G -Mixup in most cases, as shown by the results in Tables 7.1
and 7.2.

E .8 A U G M E N TAT I O N S T R AT E G I E S

We present several graph data augmentation strategies, each defined by the augmentation function Aλ .
For simplicity, we denote it Aλ (Gn ) instead of the more explicit Aλ (Gn , yn ).
Edge Perturbation. Randomly adding or removing edges. Given a graph Gn = (Vn , En , Xn ), edge
perturbation is defined as Aλ (Gn ) = (Vn , En ∪ Eadd \ Eremove , Xn ), where edges in Eadd and Eremove are
sampled from a Bernouli distribution P (λ) = B(λ) with probability λ.
Node Feature Perturbation. Augmenting node features by introducing noise or masking some features.
This is given by Aλ (Gn ) = (Vn , En , Xn + λZ), where Z ∼ N (0, I ) for Gaussian noise addition.
E .9 G R A P H D I S TA N C E M E T R I C S 185

Table e.6 – Classification accuracy (± std) on different benchmark datasets for the data augmentation baselines
using the GIN backbone. Higher accuracy (in %) is better.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

DropEdge 71.70±3.25 43.67±6.58 72.36±10.40 66.31±5.49 70.20±2.85


DropNode 71.00±4.49 42.67±5.54 75.00±6.67 64.60±4.52 70.80±3.31
SubMix 71.10±4.15 42.80±3.93 84.53±6.66 60.10±3.78 70.03±2.61

Subgraph Sampling. Extracting subgraphs. A common approach is k-hop neighborhood sampling,


where Aλ (Gn ) = (Vn′ , En′ , X′n ), with Vn′ ⊆ Vn and En′ being edges induced by the k-hop neighbors of
randomly selected nodes selected based on a prior distribution P (λ).
GRATIN. In our approach, the augmented hidden representations hGe = Aλ (hG ) corresponds to a
sampled vector from the GMM distribution Pc that was previously fit on the hidden representations
Hc = {hG | yn = c} of the graphs in the training set with the same class c. Formally, hGe = Aλc ({Hc | G ∈
c}) = λc , where λc are the parameter of the GMM distribution Pc .
It is important to note that in some augmentation strategies, the augmented graphs Genm may not explicitly
depend on the specific training graph Gn . Instead, they may be sampled or generated based on other
factors, such as a general graph distribution or global augmentation rules. This flexibility allows the
augmentation framework to capture a broader range of variations while maintaining consistency with the
original training data.

E .9 G R A P H D I S TA N C E M E T R I C S

Let us consider the graph structure space (A, ∥·∥A ) and the feature space (X, ∥·∥X ), where ∥·∥G and
∥·∥X denote the norms applied to the graph structure and features, respectively. When considering only
structural changes with fixed node features, the distance between two graphs Ge, G is defined as

Ge − G = ∥A − A
e ∥G , (e.3)

e A are respectively the adjacency matrix of Ge, G , and the norm ∥·∥G can be for example the
where A,
Frobenius or spectral norm. If both structural and feature changes are considered, the distance extends to:

Ge − G = α∥ A − A
e∥A + β∥X − X
e ∥X , (e.4)

e X are the node feature matrices of Ge, G respectively, and α, β are hyperparameters controlling the
where X,
contribution of structural and feature differences.
In most baseline graph augmentation techniques, such as G -Mixup, SubMix, and DropNode, the
alignment between nodes in the original graph G and the augmented graph Ge is known. However, in
cases where the node alignment is unknown, we must take into account node permutations. The distance
between the two graphs is then defined as
 
Ge − G = min α∥ A − P AP e T ∥A + β ∥ X − P X
e ∥X , (e.5)
P∈Π

where Π is the set of permutation matrices. The matrix P corresponds to a permutation matrix used to
order nodes from different graphs. By using Optimal Transport, we find the minimum distance over the
set of permutation matrices, which corresponds to the optimal matching between nodes in the two graphs.
186 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

This formulation represents the general case of graph distance, which has been used in the literature
(Abbahaddou et al., 2024a).
A common choice for measuring the distance between two graphs is the norm applied
q to their adjacency
matrices. One widely used norm is the Frobenius norm, defined as ∥ A − A∥ F = ∑i,j ( Aij − A
e eij )2 , which
captures element-wise differences between the adjacency matrices. Another commonly used norm is the
spectral norm, defined as ∥ A − A
e∥2 = σmax ( A − A
e), where σmax denotes the largest singular value of the
difference matrix.

E .10 G AU S S I A N M I X T U R E M O D E L S

GMMs are probabilistic models used for modeling complex data by representing them as a mixture of
multiple Gaussian distributions. The probability density function p(x) of a data point x in a GMM with K
Gaussian components is given by:
K
p(x) := ∑ πk N (x | µk , Σk ), (e.6)
k =1

where πk is the weight of the k-th Gaussian component, with πk ≥ 0 and ∑kK=1 πk = 1, and N (x | µk , Σk ) is
the Gaussian probability density function for the k-th component, defined as,
 
1 1 ⊤ −1
N (x | µk , Σk ) := × exp − (x − µk ) Σk (x − µk ) ,
(2π )d/2 det(Σk )1/2 2

where µk and Σk are respectively the mean vectors and the covariance vectors of the k-th Gaussian
component, and d the dimensionality of x. The parameters of a GMM are typically estimated using the
EM algorithm (Dempster, Laird, and Rubin, 1977), which alternates between estimating the membership
probabilities of data points for each Gaussian component (Expectation step) and updating the parameters
of the Gaussian distributions (Maximization step). GMMs are a powerful tool in statistics and machine
learning and are used for various purposes, including clustering and density estimation (Naim and Gildea,
2012; Ozertem and Erdogmus, 2011; Zhang et al., 2021b).

E .11 E X P E R I M E N TA L S E T U P

Datasets. We evaluate our model on five widely used datasets from the GNN literature, specifically
IMDB-BIN, IMDB-MUL, PROTEINS, MUTAG, and DD, all sourced from the TUD Benchmark (Morris et al.,
2020b). These datasets consist of either molecular or social graphs. Detailed statistics for each dataset are
provided in Table e.7. We split the dataset into train/test/validation set by 80%/10%/10% and use 10-fold
cross-validation for evaluation following the recent work of Zeng et al. (2024). If a dataset does not contain
node features, we follow the standard practice in GNN literature by using one-hot encoding of node degrees
as input features.
Baselines. We benchmark the performance of our approach against the state-of-the-art graph data
augmentation strategies. In particular, we consider the DropNode (You et al., 2020), DropEdge (Rong et al.,
2019), SubMix (Yoo, Shim, and Kang, 2022), G -Mixup (Han et al., 2022) and GeoMix (Zeng et al., 2024).
Implementation Details. We used the PyTorch Geometric (PyG) open-source library, licensed under
MIT (Fey and Lenssen, 2019). The experiments were conducted on an RTX A6000 GPU. For the datasets
from the TUD Benchmark, we used a size base split. We utilized two GNN architectures, GIN and GCN,
both consisting of two layers with a hidden dimension of 32. The GNN was trained on graph classification
tasks for 300 epochs with a learning rate of 10−2 using the Adam optimizer Kingma and Ba, 2015b. To
E .12 F I S H E R - G U I D E D G M M A U G M E N TAT I O N 187

Table e.7 – Statistics of the graph classification datasets used in our experiments.

Dataset #Graphs #Features Avg. Nodes Avg. Edges #Classes

IMDB-BIN 1,000 - 19.77 96.53 2


IMDB-MUL 1,500 - 13.00 65.94 3
MUTAG 188 7 17.93 19.79 2
PROTEINS 1,113 3 39.06 72.82 2
DD 1,178 82 284.32 715.66 2

model the graph representations of each class, we fit a GMM using the EM algorithm, running for 100
iterations or until the average lower bound gain dropped below 10−3 . The number of Gaussians used in the
GMM is provided in Table e.8. After generating new graph representations from each GMM, we fine-tuned
the post-readout function for 100 epochs, maintaining the same learning rate of 10−2 .
Computation of Influence Scores. Computing and inverting the Hessian matrix of the empirical
risk is computationally expensive, with a complexity of O( N × p2 + p3 ), where p = |θ | is the number of
parameters in the GNN. To mitigate the cost of explicitly calculating the Hessian matrix, we employ implicit
Hessian-vector products (iHVPs), following the approach outlined in Koh and Liang (2017).

Table e.8 – The optimal number of Gaussian distributions in the GMM for each pair of dataset and GNN backbone.

Model IMDB-BIN IMDB-MUL MUTAG PROTEINS DD

GCN 40 50 10 10 2
GIN 50 5 2 2 50

E .12 F I S H E R - G U I D E D G M M A U G M E N TAT I O N

Figure e.1 illustrates the impact of removing augmented representations on test accuracy in the Fisher-
Guided GMM Augmentation experiment. At the beginning, removing augmented graphs with low or negative
influence scores improves generalization. The highest test accuracy is reached when a significant portion
of low-quality augmentations has been removed while retaining high-influence ones. This indicates that a
well-selected augmentation subset enhances model performance. As more augmentations are removed,
the overall diversity of the training set decreases. Since data augmentation generally helps the model
generalize better, excessive removal reduces its effectiveness. At 100% removal, augmentation is entirely
disabled, meaning the model is trained only on the original dataset, i.e., the reference case, leading to a
significant drop in accuracy.

E .13 SOFTMAX CONFIDENCE AND ENTROPY DISTRIBUTIONS

One of the critical challenges in training GNN is Softmax saturation, where the model produces confident
predictions. This high confidence leads to vanishing gradients, making the influence score of the augmented
graphs converge to zero. To analyze this phenomenon, we examine the Softmax confidence and entropy
188 A P P E N D I X : I M P R O V I N G G E N E R A L I Z AT I O N I N G N N S T H R O U G H D ATA A U G M E N TAT I O N

PROTEINS Dataset
0.74
0.72

Test Accuracy
0.70
0.68
0.66
0.64
0.0 0.2 0.4 0.6 0.8 1.0
Percentage of removed augmented representations
Figure e.1 – Effect of Filtering Augmented Representations on Test Accuracy

distributions for GCN and GIN across the different graph classification datasets. In Figure e.2, e.3, e.4, e.5
and e.6, we present histograms illustrating the distribution of maximum Softmax confidence and entropy
for models trained on the original DD, IMDB-BIN, IMDB-MUL, MUTAG, and PROTEINS datasets. Each
dataset panel contains two histograms, the Confidence Histogram and the Entropy Distribution. The
Confidence Histogram illustrates the distribution of the maximum Softmax confidence scores assigned
by the model to the predicted class. A higher confidence value indicates that the model is more certain
about its classification decision. The Entropy Distribution provides a measure of uncertainty in the model’s
predictions, computed as: H (y) = − ∑c yc log(yc ), where, for each class c, yc is the predicted probability
corresponding to this class. Lower entropy values reflect high-confidence predictions and higher entropy
values indicate more significant uncertainty.
Specifically, in the DD dataset, we observe an extreme case of Softmax saturation in GIN, where almost
all predictions collapse to near-maximum confidence. The entropy histogram further reinforces the Softmax
saturation in GIN on DD, where entropy values are heavily skewed towards zero, meaning the model rarely
assigns significant probability mass outside the predicted class. Compared to other dataset-model settings,
this effect is particularly pronounced in GIN trained on DD. In contrast, GCN exhibits a wider confidence
distribution, maintaining a more balanced uncertainty.

IMDB-BIN Dataset
Confidence Histogram Entropy Distribution
10 GIN 12 GIN
8 GCN 10 GCN
8
Density

Density

6
6
4
4
2 2
0 0
0.5 0.6 0.7 0.8 0.9 1.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7
Maximum Softmax Confidence Entropy

Figure e.2 – Softmax Confidence and Entropy Distributions for the IMDB-BIN Dataset.
E .13 S O F T M A X C O N F I D E N C E A N D E N T R O P Y D I S T R I B U T I O N S 189

IMDB-MUL Dataset
Confidence Histogram Entropy Distribution
8 GIN 5 GIN
7 GCN GCN
6 4
Density

Density
5 3
4
3 2
2 1
1
0 0
0.4 0.5 0.6 0.7 0.8 0.9 1.0 0.0 0.2 0.4 0.6 0.8 1.0
Maximum Softmax Confidence Entropy

Figure e.3 – Softmax Confidence and Entropy Distributions for the IMDB-MUL Dataset.

MUTAG Dataset
Confidence Histogram Entropy Distribution
20.0 GIN 6 GIN
17.5 GCN 5 GCN
15.0
Density

Density
12.5 4
10.0 3
7.5 2
5.0
2.5 1
0.0 0
0.5 0.6 0.7 0.8 0.9 1.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7
Maximum Softmax Confidence Entropy

Figure e.4 – Softmax Confidence and Entropy Distributions for the MUTAG Dataset.

PROTEINS Dataset
5
Confidence Histogram 14
Entropy Distribution
GIN GIN
4 GCN 12 GCN
10
Density

Density

3 8
2 6
4
1 2
0 0
0.5 0.6 0.7 0.8 0.9 1.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7
Maximum Softmax Confidence Entropy

Figure e.5 – Softmax Confidence and Entropy Distributions for the PROTEINS Dataset.

DD Dataset
120
Confidence Histogram Entropy Distribution
GIN 35 GIN
100 GCN 30 GCN
80 25
Density

Density

60 20
40 15
10
20 5
0 0
0.5 0.6 0.7 0.8 0.9 1.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7
Maximum Softmax Confidence Entropy

Figure e.6 – Softmax Confidence and Entropy Distributions for the DD Dataset.

Vous aimerez peut-être aussi