Legaltechdaily iconLegaltechdailySep 9, 2026 ~8 min source read

How ontologies and a context engine stop AI guesses, cut token use, and make firm data reliable

Generic models guess when they don’t understand a firm’s labels and relationships. Building a layered context engine — glossaries, ontology, firm model, schema contracts, and a knowledge graph — gives AI the firm’s language so outputs are accurate, faster, and cheaper.

Share this story

Send the public story page.

Useful takeaways from this story.

A context engine layers meaning (entities, glossaries, ontology, firm model, schema contracts, knowledge graph) so the AI consumes a unified semantic layer tied to the firm’s records.

Governed schema contracts and an operational knowledge graph reduce back-and-forth prompts, saving time and tokens while increasing confidence in outputs.

Private capital firms and investment banks often need routine summaries — for example, a report of team engagements across LPs, advisors, and portfolio companies. When a model sees a firm record labeled "interaction," it does not know whether that means a call, a due-diligence meeting, or a fundraising touchpoint. Lacking that mapping, the model guesses. The result can be plausible but incorrect outputs that require multiple correction rounds, each costing time, tokens, and trust.

This is not primarily a model failure. It is a data architecture failure: the AI was connected to the system of record without being taught the firm's language and relationships first.

# Who is most exposed

Mid-market firms typically run lean tech teams. The person connecting an assistant to a CRM or deal system may not have deep data-architecture experience. When the tool misclassifies a deal stage or relationship, there is no specialist to catch it.

Large firms have more resources but greater exposure: more systems, more users, and higher-stakes decisions. A wrong answer about fund exposure or LP commitment at scale can travel into investor decks, committee materials, or filings.

# What fixes the problem: the context engine stack

  • Entities: raw objects tracked by the firm (deals, companies, funds, contacts). Alone they are just records.
  • Glossaries: agreed business definitions so terms like "margin" or "commitment" resolve to a single meaning across users.
  • Ontology: a structured map of concepts and relationships. It defines classes, relationships, and rules (for example, that a fund holds investments, an LP commits to a fund, an interaction connects people to a deal, and certain combinations are not allowed).
  • Firm model: the ontology configured for the firm's taxonomies, deal stages, and vintages.
  • Knowledge graph: the ontology applied to actual records so entities become nodes and relationships become edges, producing a navigable network.
  • Semantic layer: the unified, governed layer the AI consumes. It encapsulates the above so the model can answer using firm-specific meanings and links rather than guessing.

# Immediate benefits

When AI queries the semantic layer instead of raw records, outputs are more accurate, require fewer prompt/response cycles, and therefore use fewer tokens and less time. Outputs also include governed context, which helps users explain numbers in meetings and reduces the risk of confident-but-wrong statements reaching external audiences.

# Practical next steps for firms

  • Inventory your entities and existing definitions. Start with the terms that are used in investor materials and compliance documents.
  • Create or formalize a glossary tied to those terms.
  • Build an ontology that maps the relationships most relevant to your workflows (funds, LPs, deals, interactions, advisors).
  • Implement schema contracts to enforce field types and permitted values across upstream systems.
  • Operationalize a knowledge graph that links real records to ontology classes and relationships.
  • Expose a single semantic layer for AI consumption so models get firm-aware context before answering.

# Bottom line

A well-built context engine teaches AI the firm's language so the model stops guessing. That reduces hallucinations, lowers token consumption through fewer correction cycles, and produces outputs people can trust in meetings, reports, and regulatory contexts.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app