LLMs Discover PDEs Using Physical Invariants as Data Input

Fan Yang, Matt Thomson· August 27, 2026 View original

Key takeaways

  • Feeding physical invariants to LLMs significantly improves PDE discovery accuracy.
  • The "data interpretation" stage triples accuracy without requiring LLM retraining.
  • LLMs can effectively act as automated theory constructors when given appropriate data representations.
  • This approach offers a practical route to accelerating scientific discovery from complex data.

Who benefits

Scientific ResearchMaterials SciencePharmaceuticalAerospaceEnergy

Summary

This research introduces "data interpretation," a stage that feeds physical invariants from spatiotemporal fields directly to Large Language Models (LLMs) for automated Partial Differential Equation (PDE) discovery. This method nearly triples the accuracy of recovered equations compared to using raw data, without requiring any training.

A major challenge in molecular sciences is bridging molecular interactions to macroscopic behaviors, a task where traditional theory building struggles to keep pace with vast modern datasets. Large Language Models (LLMs) offer a promising avenue for automating theory construction, but directly inputting raw spatiotemporal field data into an LLM prompt is not effective. Current LLMs typically learn about data only through a score indicating how well a proposed equation fits. This paper introduces a novel "data interpretation" stage. Instead of raw data, this stage measures the field into physically meaningful quantities, such as invariants, that a human theorist would consult. These interpreted quantities are then supplied directly to the LLM as input. On a benchmark of simulated fields, this interpretation method dramatically improves the accuracy of discovered Partial Differential Equations (PDEs), nearly tripling it compared to feeding raw data, all without any additional model training. This approach enables LLMs to "read" field data in a manner analogous to a human theorist, opening a practical path for automated field theory construction that can evolve alongside experimental data.

Why it matters

For professionals in scientific research, engineering, and materials design, this method offers a powerful new tool for accelerating the discovery of fundamental physical laws and models from complex experimental data, potentially revolutionizing scientific inquiry.

How to implement this in your domain

  1. 1Explore integrating "data interpretation" modules into existing scientific discovery pipelines that use LLMs.
  2. 2Apply this technique to analyze complex experimental data in physics, chemistry, or biology to discover underlying PDEs.
  3. 3Develop tools to automatically extract physical invariants from spatiotemporal datasets for LLM input.
  4. 4Collaborate with domain experts to identify relevant physical invariants for specific scientific problems.
  5. 5Train research scientists on leveraging LLMs for automated theory construction with interpreted data.

Original post by Fan Yang, Matt Thomson

"arXiv:2608.25189v1 Announce Type: new Abstract: Understanding how molecular interactions govern macroscopic behaviour is a central challenge in molecular sciences. However, conventional theory building cannot keep pace with the vast datasets modern experimentation routinely produ…"

View on X

Originally posted by Fan Yang, Matt Thomson on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI Engineering & DevToolsAI Research

Resilient Decentralized Federated Learning for Wireless IoT Networks

This paper introduces QEF-GT-AdamW, a communication-efficient and outage-resilient algorithm for decentralized federated learning over wireless IoT networks. It combines gradient tracking, AdamW optimization, and dual-stream biased quantization with error feedback to improve robustness and convergence under heterogeneous data and unreliable communication.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI Engineering & DevToolsAI Research

FedQoS Predicts QoS Risk for Wireless Access Selection

This paper proposes FedQoS, a federated QoS-risk learning framework that predicts future QoS degradation for reliable access selection in heterogeneous indoor-outdoor wireless environments. It enables access nodes to locally learn from network logs and collaboratively train a global predictor without centralizing user data, significantly reducing QoS failure rates.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Zerihun Huruy, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI ResearchAI Engineering & DevTools

Parametric Knowledge Graphs Show Storage-Retrieval Gap

This paper explores compiling knowledge graphs into LoRA adapters for parametric memory, finding that while adapters effectively store factual knowledge, retrieving it via semantic similarity or weight-space geometry is ineffective. This highlights a "storage-retrieval gap" and the need for new query-conditioned composition mechanisms.

Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker TrespAug 27, 2026