Generic AI has knowledge. Entity management requires judgment

[email protected]

August 19, 20267 min read

Give any LLM a document, and it can easily summarize governing provisions, draft a resolution, extract key facts, or answer an entity management question in seconds. That’s one reason they’ve been so widely adopted by private market firms. 

But while LLMs are useful for lower-stakes, individual productivity tasks, they lack the judgment needed for high-stakes, complex private market workflows. 

The distinction was central to our recent webinar, Closing the Entity Management Gap with Purpose-Built AI, featuring Adrienne Williams, Product Manager at Ontra, and Jill Hyde, Deputy General Counsel and Managing Director at Beacon Capital Partners.

The argument was not that generic AI has no place in the private markets. Beacon employees already use Copilot for routine tasks such as drafting emails, and Hyde sees clear potential for AI to help teams search organizational documents, locate entity information, and answer questions faster. The point is that usefulness is not the same as judgment.

The tension arises when firms confuse the ability to produce an answer with the ability to determine whether that answer is safe to act on. LLMs are trained on broad knowledge, but they do not inherently possess the context, controls, or institutional judgment of your firm. Which of these three conflicting spreadsheets is authoritative? Has this operating agreement been superseded? Should this seemingly routine inconsistency be ignored, investigated, or escalated?

Generic AI does the best it can with the information it’s given, but in high-stakes workflows like entity management, “mostly right” is a huge risk masquerading as efficiency. 

A model can know the rule and still reach the wrong answer

LLMs are trained to recognize patterns across enormous bodies of information. They can explain how an entity is typically formed, how a board resolution is generally structured, or which requirements commonly apply in a jurisdiction.

Unfortunately, legal and compliance workflows do not operate in the world of “typically,” “generally,” or “commonly.” They need to know what applies to this entity, under these governing documents, based on the current ownership structure, subject to this firm’s policies and approval requirements. That is the gap between knowledge and judgment.

This isn’t a hypothetical risk. Stanford researchers tested ChatGPT 3.5, PaLM 2, and Llama 2 by asking specific, verifiable questions about randomly selected federal court cases; the models hallucinated between 69% and 88% of the time. The researchers also found the models were confident in their answers regardless of whether those answers were actually correct. Microsoft’s own documentation on this exact problem draws the same line: “LLMs can reason about wide-ranging topics, but their knowledge is limited to the public data that was available at the time they were trained.” For private data, Microsoft says organizations need retrieval-augmented generation: retrieving the relevant information first, then inserting it into the model’s prompt. Even then, Microsoft maintains a dedicated “groundedness detection” feature to catch outputs that still don’t match the provided source material,a tacit admission that grounding a model in the right data doesn’t automatically guarantee it uses it correctly. The source behind an answer matters as much as the model generating it.

“Places where we don’t have a true or reliable source of data are where it doesn’t work well,” Hyde noted. At Beacon, teams had historically stored information in different places and, therefore,  pointed AI at different source materials. Predictably, they received different answers. AI tools performed much better when they could work from records the firm knew were accurate, consistently maintained, and organized.

The demo starts where the real work ends

Generic AI looks most impressive under controlled conditions: when the correct document has already been identified, the relevant entity is clear, the ownership information is up to date, and the prompt contains all the important details, with no conflicting records, missing fields, or obscure exceptions.

Under these stringent conditions, any major model can produce a convincing result. Unfortunately, creating those ideal conditions is often a tremendous amount of work for private market firms. 

Entity data may be spread across shared drives, outside counsel’s records, inboxes, spreadsheets, accounting systems, and slide decks. The legal team may maintain the organizational documents while finance holds another set of records, and deal professionals retain copies from a prior transaction.

At Beacon, Hyde described a lean legal department responsible for creating and filing documents and helping colleagues across the country find them. “We’re the place where it all starts,” she said, adding that helping people locate the documents they need is one of the team’s biggest bottlenecks.

That fragmentation shows up in small, specific ways. At one growth equity firm managing a bulk entity data upload, the AI skipped a registration ID for several entities, not because the model failed, but because that field simply wasn’t in the source spreadsheet it was working from.

Generic AI does not eliminate fragmentation. In some cases, it makes it harder to see. A user can point a model at an outdated spreadsheet and receive a polished answer. Another user can point the same model at a different spreadsheet and receive another polished answer. Both outputs may sound authoritative, but neither model knows which source the firm intended to control.

“If the ownership data is stale or the jurisdictions are mistaken, AI will still work, but it will execute confidently on bad information, which is worse than a human doing it in the first place,” Williams put it plainly. 

That is the uncomfortable reality behind many AI pilots: failure does not always look like failure. The model does not refuse the task, expose the contradiction, or flag that a critical record may be missing; it simply finishes the work. One legal lead at a multi-strategy investment firm found this out mid-pilot: an AI tool meant to log sideletter provisions kept repeating the same errors a person would have caught, to the point where she was better off logging the terms herself than correcting the AI’s mistakes one by one.

A fast answer is not valuable if no one can defend it

The industry is already running into this trust gap. In KPMG’s Q4 2025 survey of asset management and private equity leaders, 53% said low trust in the accuracy and fairness of AI outputs was the biggest challenge to demonstrating AI-related ROI, the single most-cited obstacle in the survey. The same survey found 68% of leaders were piloting AI agents, while only 24% had moved to actually deploying them.

Regulators are moving in the same direction. The SEC’s 2026 examination priorities specifically flag firms’ use of AI, data sources, and “AI washing” for review. Examiners plan to assess whether firms have designed and implemented policies to monitor and supervise their use of AI technologies across back-office operations, fraud prevention, and trading functions. An AI system that can’t show its work is an exam finding waiting to happen.

Firms are struggling with generic AI because output alone is not valuable. An answer becomes valuable when a user can determine:

  • Which records informed it
  • Whether those records are current and authoritative
  • Which entity and jurisdiction does it apply to
  • Whether an exception or inconsistency was identified
  • Which firm policy governed the workflow
  • Who reviewed and approved the resulting action

That is what judgment looks like operationally. It is the ability to apply relevant context, distinguish authority from information, recognize meaningful exceptions, and route consequential decisions to the right person. Generic AI does not come out of the box with that institutional context built in; it requires constant maintenance from your team. 

Purpose-built AI is not generic AI with a legal prompt.

The difference between generic and purpose-built AI is sometimes reduced to specialization: one model knows a little about everything, while the other knows more about a particular industry. The real difference is whether the system helps users make defensible decisions. 

Purpose-built AI should not simply know more legal terminology. It should operate inside a system designed around the decisions users need to make.

In entity management, that means grounding AI in structured entity records, ownership relationships, governing documents, jurisdictional information, permissions, and defined workflows. It means showing users where an answer came from, maintaining a record of changes, and introducing human review where accuracy or risk demands it. Ontra’s Atlas solution does this. 

Upload a structure chart, and Atlas can extract entities, relationships, and ownership percentages into structured data, but that extraction doesn’t automatically become the record of truth. The user reviews what the AI found, corrects errors, adds missing information, and confirms the structure before it’s saved. Questions get answered the same way: Atlas’s Ask AI feature draws on the entity information and relationships already stored in the platform, and links its answers back to the underlying entity records rather than returning an unsupported block of text.

As of the end of 2025, Atlas covered 185 jurisdictions and 30 entity types across 68 customers, spanning 29,310 entities created on the platform. It’s the kind of reconcilable, firm-specific dataset a generic model doesn’t have on its own, regardless of how capable the underlying LLM is.

None of this is an argument against using AI in entity management. It’s an argument against confusing a fast answer with a safe one. Generic models are good at producing something that sounds right; the question worth asking of any AI system is whether it can also show which record an answer came from, whether that record is the one the firm actually trusts, and who signed off before it became fact. Private equity firms don’t just need AI that knows more; they need AI whose answers they can hold accountable.

 

Explore Category

Explore Tags