LLMs Enhance Data Lake Metadata for Relationship Discovery

Ahlame Diouan (ERIC, UL2), Eric Ferey (ERIC, UL2), Sabine Loudcher (ERIC, UL2), J\'er\^ome Darmont (ERIC, UL2)· August 28, 2026 View original

Key takeaways

  • LLMs can significantly improve metadata quality and relationship discovery in data lakes.
  • ColRel uses a two-stage method, including business dictionaries, to interpret complex schema labels.
  • The approach is particularly effective for ERP-derived datasets with weak semantic signals.
  • Automating metadata enrichment can reduce manual effort and accelerate data integration.

Who benefits

Data ManagementEnterprise SoftwareAnalyticsFinanceHealthcare

Summary

This paper introduces ColRel, a two-stage method that uses large language models and business dictionaries to build column embeddings from metadata and data, improving relationship discovery in data lakes, especially for complex ERP datasets. Experiments show its effectiveness in semantically related, weak-signal settings.

Data lakes often suffer from limited or uninformative metadata, making it difficult to identify relationships between data columns, particularly in complex enterprise resource planning (ERP) systems with coded schema labels. A new method, ColRel, addresses this by employing a two-stage process. It first generates column embeddings using available metadata and data during ingestion. For challenging cases involving coded schemas, ColRel leverages business dictionaries to better interpret column names. These interpretations are then used to create short natural-language descriptions, which feed into the second stage of the embedding process. This approach significantly improves the discovery of semantic relationships, even in scenarios with weak signals.

Why it matters

Professionals managing large data estates can leverage this approach to improve data discoverability and integration, reducing manual effort in understanding complex datasets and accelerating data-driven initiatives.

How to implement this in your domain

  1. 1Evaluate current data lake metadata quality and identify gaps in column relationship discovery.
  2. 2Explore integrating business dictionaries or glossaries to enrich existing metadata.
  3. 3Pilot a two-stage embedding process using LLMs to generate semantic descriptions for data columns.
  4. 4Develop automated tools to apply this method at data ingestion time for new datasets.
  5. 5Measure the improvement in data discoverability and the efficiency of data integration projects.

Original post by Ahlame Diouan (ERIC, UL2), Eric Ferey (ERIC, UL2), Sabine Loudcher (ERIC, UL2), J\'er\^ome Darmont (ERIC, UL2)

"arXiv:2608.26750v1 Announce Type: new Abstract: Data lakes rely on metadata to remain usable, yet this meta data is often limited or weakly informative for column relationship discovery, especially in ERP-derived datasets with coded or abbreviated schema labels. We propose ColRel…"

View on X

Originally posted by Ahlame Diouan (ERIC, UL2), Eric Ferey (ERIC, UL2), Sabine Loudcher (ERIC, UL2), J\'er\^ome Darmont (ERIC, UL2) on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools