Stochastic Knowledge Graphs Improve LLM Student Simulation

Yuan An, Emily Wang, Benjamin Wang, Ruhma Hashmi· August 25, 2026 View original

Key takeaways

  • Traditional LLM student simulations struggle to accurately represent low mastery levels.
  • Stochastic Student Knowledge Graphs (SSKG) provide a more faithful simulation method.
  • SSKG uses knowledge graphs and probabilistic sampling to determine answer correctness.
  • This approach generates realistic mastery gradients, improving educational AI development.

Who benefits

EdTechEducationAI DevelopmentLearning & Development

Summary

This paper introduces Stochastic Student Knowledge Graphs (SSKG) to create more faithful LLM student simulations, addressing the limitation of prompt-based methods where LLMs struggle to accurately simulate low mastery. SSKG significantly reduces accuracy and produces a clear mastery gradient, making simulations more realistic.

Large Language Models (LLMs) are increasingly used to simulate students at various mastery levels, which is valuable for generating synthetic training data and stress-testing tutoring systems. However, existing prompt-based simulation methods often fall short because LLMs tend to perform at their inherent high capability, even when instructed to simulate a student with low mastery. This makes it difficult to differentiate between low and high mastery profiles. Researchers demonstrated this limitation using 379 College Board-calibrated SAT Algebra items, where three leading LLMs achieved near-perfect accuracy (96.8-100%) across all simulated mastery profiles. To overcome this, the study proposes a novel method grounded in a Stochastic Student Knowledge Graph (SSKG). This approach begins by extracting a curriculum knowledge graph (CKG) from an open algebra textbook, then decomposes each SAT solution into a chain of required knowledge triples. The SSKG assigns a mastery probability to each knowledge triple, which is then sampled to determine the correctness of a student's answer. An LLM subsequently generates a first-person rationale consistent with this sampled outcome. This SSKG-based simulation significantly reduced accuracy to a more realistic range of 44.1-85.2% across different mastery profiles and, crucially, produced a clear, monotonic mastery gradient. This method offers a more faithful and nuanced way to simulate student performance, providing better data for educational AI development.

Why it matters

For professionals developing educational AI, tutoring systems, or adaptive learning platforms, more faithful student simulations enable better testing, more realistic data generation, and ultimately, the creation of more effective and personalized learning experiences.

How to implement this in your domain

  1. 1Adopt Stochastic Student Knowledge Graphs (SSKG) for generating synthetic student data to train and evaluate educational AI systems.
  2. 2Integrate knowledge graph extraction and probabilistic sampling into LLM-based student simulators to improve fidelity.
  3. 3Benchmark existing LLM student simulation methods against SSKG to identify areas for improvement in mastery differentiation.
  4. 4Utilize more realistic student simulations to stress-test adaptive learning algorithms and personalized tutoring systems.

Original post by Yuan An, Emily Wang, Benjamin Wang, Ruhma Hashmi

"arXiv:2608.21668v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches le…"

View on X

Originally posted by Yuan An, Emily Wang, Benjamin Wang, Ruhma Hashmi on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses