AI for legal and compliance in PE: the ground truth problem no model can solve

Eric Hawkins

September 2, 20267 min read

Why AI adoption outpaced AI trust in private markets

The mandate inside private market firms has been simple: adopt AI wherever you can, fast, or fall behind. And it worked. AI adoption in private markets deal workflows more than doubled between 2024 and 2025, with nearly a third of firms now using it for due diligence. Major model improvements reduced the cost of building in-house, and the easiest work—retrieval, first-pass drafting, summarizing a document—really is faster now.

But the industry is now hitting the trough of disillusionment. After a few months of using generic AI tools, the magic has worn off, and firms are stuck asking the harder question: how do I know when it’s wrong?

Engineering teams have been wrestling with AI-generated work for years because writing software was the first real use case for these models. We learned early that you cannot hand an agent a task and trust the result until you’ve built something that checks its work at scale—automated tests, verification frameworks—all the infrastructure that tells you if an AI-written program actually does what it says. Building that operational muscle came before the automation, not after, with decades of software development best practices to light the way.

Ground truth: It's the only way you check whether an AI system got something right at scale, instead of eyeballing every answer yourself.

Eric Hawkins

 | CTO at Ontra

Ground truth: the record that checks the model

In software engineering, automated tests are used to assert that AI-generated code does the right thing. In AI engineering, the record of what a correct answer looks like—built from expert humans making the same call over and over—has a name: ground truth. It’s the only way you check whether an AI system got something right at scale, instead of eyeballing every answer yourself.

It’s also a way of defining judgment in legal and compliance. Every negotiation, every edge case, every near-miss you caught and corrected, that’s your own version of ground truth.

At a small scale, none of this matters. If you’re the GC or the CCO with ten side letters you know by heart, asking an AI tool a question and eyeballing the answer is completely fine. But once you’re managing thousands of side letters across strategies you didn’t personally negotiate, with a dozen other attorneys touching the same documents, there’s no eyeballing your way to confidence. You need ground truth at scale, and almost no firm has it because it’s not something they’ve ever had to codify.

An investment firm we work with found this out the hard way. Their in-house build didn’t hallucinate; it overshared, including information in a written response that was never meant to be disclosed. That’s not a training-data problem you can fix with a better prompt. It’s what happens when a tool that’s excellent at retrieval lacks ground truth about what should or shouldn’t be said in front of an LP or a regulator. Science fiction has spent a century imagining that, with enough rules, you could teach robots to know right from wrong. The genre’s whole anxiety is that judgment doesn’t work that way.

Neither, it turns out, does a side letter. Legal and compliance teams were asked to skip straight to AI-driven automation, with no ground truth to validate the outputs.

As a result, only 33% of legal professionals trust AI-assisted legal work results. 67% worry the verification cost outweighs AI’s efficiency gains. On another call about trigger-based obligations—the kind tied to an event with no calendar date—one investor at an asset management firm said it himself: “We know things the system can’t possibly know. Someone has to tell it.”

Where generic AI works and where it structurally can’t

None of this means AI is failing private market firms broadly. I’ve seen teams get incredible value in deal diligence, underwriting, and valuation. Deal analysis rewards a model that can synthesize a lot of information into something directionally useful. Those teams already had ground truth sitting around—years of IC memos, valuation models, and underwriting notes—the encoded judgment from the firm’s own history.

A plausible-sounding wrong answer is worse than no answer at all

Eric Hawkins

 | CTO at Ontra

Legal and compliance work is structurally different. It’s binary. There’s never a roughly correct answer to “am I allowed to do this?” or “did we send this required notice?” A plausible-sounding wrong answer is worse than no answer at all, and without ground truth to check it against, you’re pointing a probabilistic tool at a problem that only accepts one right answer.

Build vs. buy: why in-house AI stalls on messy data

I wrote before about the three signals that tell you building in-house is the wrong call: cross-team operational complexity, a high cost of error, and institutional knowledge that isn’t written down anywhere.

I’d add that the single biggest reason in-house builds struggle is that the firm’s own data is a mess, and almost every CIO or CTO I talk to admits it in private, even if they won’t say it out loud in a board meeting. Side letters are scattered across SharePoint with no record of which version supersedes which. Every team has different strategies for storing documents. Outside counsel is holding pieces that nobody centralized. Point a model at that, and it will give you a very confident, very wrong answer.

The average private markets firm is, in this sense, a fiefdom: different strategies, different funds, different LPs, each with its own context that doesn’t automatically travel to the next room. You can’t fix that by pointing a better model at it. You fix it by doing the unglamorous work of mapping that context and building the ground truth first, which is exactly the work most in-house teams don’t have the time, headcount, or appetite to do themselves.

That’s the part that generic AI cannot give you, no matter how good the underlying model gets. Robert Smith, founder and CEO of Vista Equity Partners, has put a version of this argument in blunter terms on Bain’s Dry Powder podcast: own the data and the workflows that define how your firm makes decisions, and buy the coordination layer someone else is already running at scale.

Where experience compounds

The reason a purpose-built system behaves differently is that someone else did the tedious work of building ground truth at a scale no single firm could justify on its own, and that memory doesn’t reset at each stage of the fund’s lifecycle.

Ontra has processed more than 2 million documents structured around private markets obligations, NDAs, side letters, and entity data. More than 53,000 side letters have been digitized with AI across Ontra’s Insight platform. But volume alone doesn’t equate to ground truth. It’s the people reviewing that volume, who’ve seen the edge case before and recorded the correct call, who help turn the AI pattern-matching into something closer to institutional memory.

AI-native services: finished work, not another output to review

Everything up to this point has been about ground truth: the record you need to check a model’s work, and the test harness that catches it when the model’s wrong. But ground truth doesn’t just appear. Experts have to curate it; catching the exception, correcting the draft, teaching the system what it missed, one matter at a time. The right tools, built intentionally, can capture those correction signals and turn them into ground truth over time. That’s real progress over a generic model with no memory.

But it still requires people to be the ones sitting inside that tool, doing the work, catching what the model misses, day after day. McKinsey’s Ben Ellencweig, who leads AI work with private equity firms and their portfolio companies, said it well on McKinsey’s Deal Volume podcast: “We need to look at the outputs of gen AI with a critical eye and apply our human judgment to evaluate whether we trust them.” Most firms simply don’t have the time or manpower to do that, so they turn to outside service providers.

And I believe that service providers should do more than hand you yet another tool. They should use those tools, on top of the ground truth they’ve already spent years curating, and hand you the finished, correct answer instead of just another output to review.

That’s what AI-native services for legal and compliance workflows mean. Fully completed work, with someone besides the model on the hook for whether it’s right. Technology that retrieves, paired with people who’ve done this before and are focused on the exceptions and judgment calls a model can’t make for itself. Buy that combination, and you get outcomes AI alone doesn’t produce: a defensible audit trail when a regulator asks, and institutional knowledge that compounds instead of leaving with whoever quits next.

The workflows where a plausible-but-wrong answer is worse than no answer—obligation tracking, side letter interpretation, anything that ends up in front of an LP or a regulator—require a kind of experience no model can generate on its own. It only accumulates the way it always has: by people being present, again and again, for the version of the problem that doesn’t look like the last one.

Explore Category

Ontra is not a law firm and does not provide any legal services, legal advice, or referral services and, as a result, we do not provide any legal representation to clients, nor do we participate in any legal representation of clients. The contents of this article are for informational purposes only, and are not intended to constitute or be relied upon as legal, tax, accounting, regulatory, or other professional advice, opinion, or recommendation by Ontra or its affiliates. For assistance or guidance regarding the impact or applicability of the topics discussed in this article to your business, please consult your legal or other professional advisers.