Ask five different legal and compliance teams which AI tool they’ve standardized on for side letter and MFN work, and you won’t get a consistent answer, often not even within the same firm. As our senior sales engineer, Ed Hawkins, summed it up during our recent webinar, Side Letters and MFN: The Work Generic AI Leaves Unfinished, “the standard is that everyone’s using everything.”
That inconsistency is more than a tooling problem; it’s a knowledge problem. A generic AI tool can read every side letter a fund has ever signed, but reading isn’t the same as knowing what any one of those provisions means once it’s actually tested in an audit, redemption, or dispute.
Nearly everyone is comfortable with the most straightforward AI use cases: drop in a handful of documents, ask a question, and get an answer, fast. It’s useful, but it only addresses the easy problem, retrieval. The hard problem is judgment at scale: your firm is being audited by the SEC, and you need to review every side letter a fund has ever signed to understand what’s been promised, to whom, and whether the fund is still keeping those promises, the kind of call a model that’s only ever read the documents, and never sat across the table when one of them mattered, has no way to make on its own.
A one-off lookup and a complete, standing record of every obligation are not the same job. Too many teams are trying to do the second one with AI tools built for the first.
Manual workflows can’t keep up
Side letter obligations used to be something one person could hold in their head. That’s no longer true. It’s also the stage of a fund’s life where the stakes compound quietly. Obligation tracking and MFN elections don’t get a single moment of scrutiny, the way a fundraise or an exit does. They run for years, through different hands, until an audit, transfer, or a wind-down forces someone to reconstruct all of it at once.
- As of July 2026, Ontra’s Insight solution manages more than 12,000 funds with over 2.5 million side letter provisions on file.
- Roughly a million side letter provisions were added in 2025 alone.
The provisions are less standardized, too. Firms that once fit MFN terms into a handful of boilerplate categories are increasingly building custom ones. It’s not just the largest LPs driving that; smaller investors are now negotiating the way only the biggest players used to.
Secondaries add another layer. Global secondary transaction volume hit $240 billion in 2025, up 48% year-over-year and the largest year on record. Every transfer is a side letter changing hands and an obligation that needs re-mapping.
Generic AI is intelligent but not necessarily complete
None of that would matter if a general-purpose AI tool scaled cleanly from five documents to five hundred. Unfortunately for most firms, it doesn’t.
Matt Crowley, GM of Product at Ontra, had a powerful metaphor: “It’s not that valuable to avoid four out of five landmines if you still step on the fifth.” If you ask a generic AI tool to surface every provision matching a query across a large batch of side letters, it may come back with four out of five. That’s cold comfort when the fifth is the obligation that surfaces in the middle of an audit. The LLM isn’t wrong, it’s just incomplete and doesn’t know it.
That gap has shown up in our conversations with fund managers. One analyst at a credit-focused fund tried using a generic AI tool to extract ESG terms from a stack of side letters. She candidly described it as ‘fast and useful’ but flagged that she had no way to confirm the output caught every relevant obligation rather than just the ones she already knew to look for. An operations lead at another large investment firm put it more bluntly: their internal AI tool would simply stall out once a batch of documents got large enough.
Hawkins made the same point about MFN elections: a process with five or ten side letters and clear-cut carve-outs can be run through a general AI tool. But push into twenty-plus carve-outs, each with its own legal interpretation, and the tool has to do real, critical thinking, exactly where, in his experience, generic AI consistently tops out.
Why not just build in-house?
If generic AI can’t be trusted to catch everything, the obvious next question is, why not build something in-house? Model access has gotten cheaper, so what’s the downside?
The GC is optimizing for legal workflows. The CCO is building compliance infrastructure. The CTO is deciding what to build next. Each call is locally rational, but together they add up to an AI stack no one owns, and no one can defend when a regulator asks how it all works. Our new guide, Building vs. Buying AI for Private Market Workflows, shows where to draw the line.
But no matter how powerful the models are, one thing that hasn’t gotten cheaper is the maintenance of those in-house systems. The permissions, auditability, and infrastructure of a purpose-built AI tool are the full-time job of someone working at a vendor or partner. At a private market firm, who will be dedicated to building and maintaining this system? Models change every few weeks, and staying on top of which model performs best for which task is strenuous, constant, and expensive to maintain in-house.
There’s also the iteration problem: a partner or provider learns from every customer and folds those lessons back into the product. “As a customer who’s building your own solution, you’re dealing with an n of one,” Crowley said. “You’re making every mistake yourself,” instead of benefiting from someone else’s, an argument our CTO Eric Hawkins echoed elsewhere in three words: “a prompt is not a repository”
None of these rules out agentic AI; it just changes where it plugs in. Crowley described customers connecting Insight’s structured, human-reviewed data to their own LLMs through our MCP, so the model still handles natural-language questions. At the same time, the answer is drawn from Insight’s dataset. Hawkins shared: a scheduled task that scans a firm’s emails and Teams messages for language matching a side letter trigger event, checks it against Insight’s structured obligations, and flags exactly which side letters it implicates. The kind of cross-referencing a spreadsheet can’t do, and a one-off chatbot query was never built for.
The completeness test
That reframes the real question for any AI tool touching obligations from “did it find the right answer” to “would it find all of them, every time, with a record of how it got there.” As Crowley put it, a human-reviewed, structured dataset is “sidestepping all five landmines. That’s something no LLM can do on its own yet.”
Hawkins had a proof point to back it up: a customer whose internal audit team was “blown over” by what a properly structured compliance calendar could do. Insight turned roughly 4,000 raw obligations into 120 affirmative ones. Noise reduced, audit trail intact, nothing missed.
None of this makes generic AI tools obsolete, nor is it an argument against AI. It’s an argument for governed AI: the same models, backed by structured data and a human record of how an answer was reached. Generic tools are still the fastest way to get a first-pass answer on a single document, and Insight’s own workflows lean on large language models throughout. The distinction is what you’re trusting the tool to do. A generic AI tool answering a one-off question about a single document is a fine use of five minutes. Trusting it to confirm, with certainty, that it caught everything across five hundred documents is a different bet, and it’s the one private markets compliance workflows can’t afford to lose.
Watch Side Letters and MFN: The Work Generic AI Leaves Unfinished, to hear Matt Crowley and Ed Hawkins walk through more of where generic AI can help and what it misses.



