The Big Picture
A two-layer graph (one that preserves every document chunk and one that encodes domain concepts) plus a router that picks the right agent gives much better, traceable answers than flat document search alone.
ON THIS PAGE
Key Findings
A lossless lexical layer keeps every document, section, table, and figure retrievable for exact lookups, while an ontology-driven domain layer collapses mentions into canonical entities to reveal cross-document patterns. Routing questions to either the lexical agent (for precise, local detail) or the domain agent (for corpus-wide aggregation and anomaly detection) improves accuracy and traceability. A 505-question benchmark shows clear gains from the lexical graph over a vector-only baseline and demonstrates where the domain graph is required. Orchestrator-Worker Pattern
Key Data
1Full graph built from 38 documents: 12,353 nodes and 45,471 edges.
2Lexical agent vs vector baseline: Tier-1 multiple-choice accuracy 95% vs 89%; Tier-2 pass rate 85% vs 65%.
33,236 document-local mentions collapsed into 537 canonical entities; 88.8% of merges done by exact or curated-synonym matching.
What This Means
Process chemists, data engineers, and AI product leads in biopharma who wrestle with years of versioned reports and handoffs—this approach makes industrial knowledge searchable, auditable, and easier to act on. Anyone building retrieval-augmented assistants or governance around them can use the dual-layer design and benchmarking protocol to know when a simple document chatbot suffices and when a domain graph is needed. Human-in-the-Loop
Test your agentsValidate against real scenarios
Key Figures

Fig 1: Figure 1: Knowledge accumulation and route convergence across a typical CMC process development timeline. Each phase draws on its own characteristic sources, named beneath it, and each node represents an information source. The most valuable insight lies not in a single record but in the relationships between them which span phases, scales, sites, and years. A platform based on knowledge graph transforms fragmented data into actionable knowledge.

Fig 2: Figure 2: The multi-layer architecture: the data layer holds raw documents in heterogeneous formats, disconnected and without shared structure. Lossless ingestion of the documents produces the knowledge layer (lexical graph), in which every document becomes its own Document → \rightarrow Section → \rightarrow Chunk tree joined by HAS_SECTION and HAS_CHUNK edges, with HAS_NEXT preserving reading; the knowledge layer is well suited to keyword and passage questions but carries no edges between documents. Entity extraction, based on a domain ontology, and global resolution produce the intelligence layer (domain graph), whose canonical entities are anchored to the chunks that mention them by HAS_MENTION edges. The intelligence layer merges lexical islands across documents and projects into one queryable graph.

Fig 4: Figure 4: (a) Tier-1 multiple-choice accuracy (green) and Tier-2 pass rate (purple), defined as a correct T1 letter, and Tier-2 judge score ≥ \geq 4, respectively, for the vector-RAG baseline and the Lexical_ReAct based on 505-question curated bank. Error bars are Wilson 95% confidence intervals. (b) T1 and T2 scores broken down by the four question cohorts. (c) Failure taxonomy over the 74 Tier-1/Tier-2 divergent items sent to Tier-3 SME review under a stratified sampling design ( Table 1 )

Fig 6: Figure 6: Question routing and the questions each layer serves. Every question enters the router agent, which invokes the domain agent (left), the lexical agent (right), or a hybrid mode holding both toolsets. The domain agent operates on process-chemistry concepts (canonical entities) and answers properties of the corpus as a whole; the lexical agent retrieves precise, locally scoped detail from chunks and their embeddings. HAS_MENTION runs from each canonical entity to the chunks that mention it, so a domain graph answer can usually be traced down to the passages that support it.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreYes, But...
The domain (ontology) layer was evaluated on a limited question set from one discontinued small-molecule program, so broader performance on multiple projects remains to be shown. Extraction steps currently show non-reproducibility between runs and need grounding/verification before regulated use. The pipeline does not read reaction-scheme images, so information embedded in figures can be missed without additional vision-to-structure tooling. Explainability
Deep Dive
The platform ingests messy CMC documents (digital, scanned, handwritten), preserves every file as a hierarchical lexical graph (Document → Section → Chunk), and builds an ontology-guided domain graph on top by extracting canonical entities and linking them back to source passages. For a testbed of 38 documents (~52,000 words), the dual-layer graph totaled 12,353 nodes and 45,471 edges. The domain layer created 537 canonical entities anchored to chunks through 22,688 mention edges, collapsing 3,236 local mentions by mostly deterministic rules (nearly 89% via exact or curated matches).
A router agent decides whether a question should be handled by a lexical agent (best for precise, locally scoped lookups) or a domain agent (best for cross-document aggregation, anomaly tracing, and process-level queries). A three-tier benchmark of 505 curated questions found the lexical graph plus agentic retrieval outperformed a vector-only baseline (95% vs 89% MCQ accuracy; 85% vs 65% pass rate) and exposed where passage retrieval fails—mainly on corpus-wide or unanchored domain questions. Limitations include extraction reproducibility, incomplete ontology depth, missing figure parsing for reaction schemes, and the need for larger domain-graph benchmarks; addressing these would make the system suitable for regulated CMC workflows and wider scientific domains. Emergence-Aware Monitoring Pattern Agent Registry Pattern
Not sure where to start?Get personalized recommendations
Credibility Assessment:
ArXiv preprint with no affiliations or citations. Slightly higher because author names may be individual researchers but still limited identifiable signals.