- AutorIn
- Shuzhou Yuan Technische Universität Dresden
- Titel
- Learning Language Models on Graphs
- Zitierfähige Url:
- https://nbn-resolving.org/urn:nbn:de:bsz:14-qucosa2-1054830
- Übersetzter Titel (DE)
- Lernen von Sprachmodellen auf Graphen
- Erstveröffentlichung
- 2026
- Datum der Einreichung
- 16.12.2025
- Datum der Verteidigung
- 19.06.2026
- Abstract (EN)
- Language models (LMs) have achieved remarkable success across natural language processing (NLP) tasks, yet they remain fundamentally limited by their reliance on sequential token representations. Natural language, however, is inherently relational and hierarchical: it reflects syntactic dependencies, discourse structures, and multi-hop reasoning chains that are not naturally encoded in linear text. Graph representations, by contrast, explicitly capture such relational structure and can be modeled effectively using Graph Neural Networks (GNNs). Despite this potential, integrating graph-based representations into LMs remains challenging due to modality differences, representational misalignment, and architectural constraints. This dissertation investigates how explicit graph-structured information can enhance language models and addresses three core challenges central to the future of LM development: architecture, efficiency, and interpretability. Because LMs primarily operate on text, we seek to bridge the gap between linguistic and graph modalities by proposing new LM architectures capable of processing graphs. As model scale increases, efficiency becomes a critical bottleneck, motivating methods for parameter-efficient training and model compression. Finally, the vast number of parameters in modern LMs complicates interpretability, demanding principled approaches for generating faithful explanations. This dissertation proposes solutions to each of these challenges across a range of NLP tasks. First, we identify key limitations of current generative LMs on graph-to-text tasks, including poor structural comprehension and susceptibility to hallucination. To address this, we introduce GraSAME, a graph-guided self-attention mechanism that incorporates token-level graph structure into pretrained LMs, enabling joint reasoning over textual and graph inputs. This architectural advance not only increases flexibility in processing graph data but also substantially improves the quality and fidelity of generated text. Second, to address efficiency challenges, we propose GNNavi, a parameter-efficient framework that uses graph neural networks to guide the internal information flow of LMs. By representing information pathways inside LMs as graphs and embedding these structures into the model, GNNavi enables effective prompt-based few-shot training for text classification while significantly reducing computational cost. We further demonstrate that many LM layers can be pruned via subgraph reduction without degrading performance, offering insights into model compression and redundancy. Third, we develop G-TEx, a graph-guided textual explanation framework that enhances the faithfulness and interpretability of natural language explanations (NLEs). By encoding model-internal highlight explanations as graph structures and injecting them into LMs through GNNs, G-TEx produces more input-grounded and faithful explanations while preserving high-quality generation. Empirical results across multiple benchmarks show that graph-based integration consistently strengthens model reasoning, reduces computational overhead, and improves the interpretability of model's reasoning process. Collectively, this work demonstrates that structural representations serve as a powerful complement to sequential token processing and provides a unified framework for improving LM architecture, efficiency, and interpretability. The findings offer actionable insights for future research, for example for molecule-structure modeling in scientific discovery and culture-aware graph-structured reasoning in social applications, and highlight the broader potential of structured information to advance next-generation, robust, and human-understandable NLP systems.
- Freie Schlagwörter (DE)
- Sprachmodelle, Verarbeitung natürlicher Sprache, Graphneuronale Netzwerke, Künstliche Intelligenz
- Freie Schlagwörter (EN)
- Language Models, Natural Language Processing, Graph Neural Network, Artificial Intelligence
- Klassifikation (DDC)
- 004
- Klassifikation (RVK)
- ST 306
- SK 890
- GutachterIn
- Prof. Dr. Michael Färber
- Prof. Dr. York Sure-Vetter
- Den akademischen Grad verleihende / prüfende Institution
- Technische Universität Dresden, Dresden
- Version / Begutachtungsstatus
- publizierte Version / Verlagsversion
- URN Qucosa
- urn:nbn:de:bsz:14-qucosa2-1054830
- Veröffentlichungsdatum Qucosa
- 06.07.2026
- Dokumenttyp
- Dissertation
- Sprache des Dokumentes
- Englisch
- Lizenz / Rechtehinweis
CC BY 4.0