English

GMLM: Bridging Graph Neural Networks and Language Models for Heterophilic Node Classification

Computation and Language 2025-10-09 v6 Artificial Intelligence Machine Learning

Abstract

Integrating Pre-trained Language Models (PLMs) with Graph Neural Networks (GNNs) remains a central challenge in text-rich heterophilic graph learning. We propose a novel integration framework that enables effective fusion between powerful pre-trained text encoders and Relational Graph Convolutional Networks (R-GCNs). Our method enhances the alignment of textual and structural representations through a bidirectional fusion mechanism and contrastive node-level optimization. To evaluate the approach, we train two variants using different PLMs: Snowflake-Embed (state-of-the-art) and GTE-base, each paired with an R-GCN backbone. Experiments on five heterophilic benchmarks demonstrate that our integration method achieves state-of-the-art results on four datasets, surpassing existing GNN and large language model-based approaches. Notably, Snowflake-Embed + R-GCN improves accuracy on the Texas dataset by over 8\% and on Wisconsin by nearly 5\%. These results highlight the effectiveness of our fusion strategy for advancing text-rich graph representation learning.

Keywords

Cite

@article{arxiv.2503.05763,
  title  = {GMLM: Bridging Graph Neural Networks and Language Models for Heterophilic Node Classification},
  author = {Aarush Sinha},
  journal= {arXiv preprint arXiv:2503.05763},
  year   = {2025}
}
R2 v1 2026-06-28T22:11:19.385Z