Citation Needed: Provenance for LLM-Built Knowledge Graphs — Daniel Chalef, Zep AI
LLM synthesis destroys data provenance trails, creating critical risks in high-stakes agent applications
“Synthesis often destroys the paper trail of how these outputs were originated.”
Daniel Chalef of Zep AI presents the provenance problem in LLM-built knowledge graphs: because LLMs synthesize facts non-deterministically from multiple sources, the audit trail of how any given fact was derived is lost. This matters acutely in regulated domains like healthcare, where an AI agent presenting a synthesized patient fact (e.g., a penicillin allergy) without source attribution could mislead clinicians. Zep's open-source Graphiti framework is their engineering response, designed to maintain traceable provenance in enterprise agent memory systems.