The Big Picture
A two-agent system can generate human-readable column descriptions and reliably tag sensitive columns across millions of enterprise fields without ever reading cell values, by fusing metadata, rules, and model reasoning to prioritize recall and auditability.
ON THIS PAGE
The Evidence
Combining three value-free signals — a metadata example retriever, rule/regex matches, and an LLM-based description signal — and fusing their ranked outputs gives better sensitivity-tag recall than any single method. The system never inspects cell values; it uses catalog keys, source code, and existing metadata so it can run on restricted datasets. Fusion improves end-to-end recall-focused performance while providing per-tag provenance and automated self-grading; the main trade-off is higher latency and cost when all strategies run. This approach aligns with the Evaluation-Driven Development pattern.
Data Highlights
1Real-world catalog scale: ~3.4 million columns across ~98,000 tables (motivating automation over manual labeling).
2Contrastive fine-tuning of the metadata encoder raised retrieval recall dramatically: NDCG@10 from 0.55 to 0.92 and MAP@100 from 0.19 to 0.90.
3End-to-end, recall-weighted F2 rose to 0.890 for the fused system versus 0.880 for the best single strategy (metadata); the fused pipeline’s median request time was 57.8 seconds.
What This Means
Platform engineers and data-platform teams who must scale catalog documentation and privacy tagging without exposing sensitive data will find this directly actionable. Privacy, compliance, and data-steering teams benefit because tags come with provenance and confidence scores, making steward review and audit easier. Machine-learning engineers building metadata services can reuse the modular fusion and self-grading patterns. For governance and audit considerations, see the Supervisor Pattern.
Not sure where to start?Get personalized recommendations
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreConsiderations
Because the system never reads cell values, some content-dependent signals are unavailable and could limit recall for tags that require value inspection. The curated evaluation is in-distribution and table-disjoint from training, but public-benchmark generalization and human–judge agreement remain to be demonstrated. Full fusion improves recall but increases latency and cost, so deployments must balance coverage versus runtime and cost. This aligns with a Hierarchical Multi-Agent Pattern.
Methodology & More
Glyph splits the problem into two cooperating agents: a Descriptor that generates authoritative, code-grounded natural-language descriptions for columns (using source code and metadata rather than cell values), and a Tagger that assigns multi-label sensitivity tags from a governed taxonomy of 275 leaf annotations. The Tagger combines three value-free strategies: a dense-retrieval over labeled metadata examples (with an encoder fine-tuned contrastively), a rule/regex matcher, and an LLM-based description tagger. Candidate tags are fused with Reciprocal Rank Fusion, validated against the ontology, and a judged LLM loop refines outputs; every tag carries provenance, score, and justification for auditability. This process can be viewed through the lens of the two cooperating agents approach, and the three value-free strategies. Evaluation uses steward-curated, per-line-of-business holdout sets and optimizes an F2 metric that weights recall more heavily (false negatives are costly for governance). Key results show strong retrieval improvements after contrastive tuning (NDCG@10 and MAP@100 jumps) and that fusion improves recall-focused F2 to 0.890 versus 0.880 from the best single strategy; regex alone is almost useless outside hand-written patterns. Operational lessons include the value of a value-free, code-grounded design (enables running on restricted tables) and batching columns per LLM call to control latency and cost. Future work suggested: collective inference across columns to exploit inter-column signals and a learned policy loop to better prioritize rare tags.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
ArXiv preprint with no affiliations or author h-index info provided. Signals point to emerging/limited information.