Vendor Due Diligence for AI-Powered CRE Software
Check the vendor's accuracy on your actual documents, not their polished demo deck.

This is about buying AI software for commercial real estate, and why the standard vendor checklist doesn't cut it. CRE documents carry fiduciary weight, decisions move real capital, and a bad output surfaces as a credit loss or a missed deal, not a support ticket. What follows lays out the specific things a firm needs to check before signing, drawn from how CRE actually works and what tends to go wrong when a tool doesn't fit that reality.
Standard IT procurement checklists were built for a different problem. They ask about uptime, integration hooks, seat licensing, maybe a SOC 2 badge somewhere in the appendix. None of that tells a buyer whether the AI reading a rent roll actually understands what a rent roll is. A bad CRM slows down a sales team's Tuesday. A bad AI underwriting tool can misprice risk on a multimillion-dollar loan, miss a covenant breach that was sitting in plain sight, or hand a credit committee a confident-sounding number that's wrong in a way nobody catches until the borrower is already underwater. That's the asymmetry this whole piece is built around, and it's why every section below treats "does it work" as a much harder question than most vendors want it to be.
The market conditions that make getting this decision right urgent
Two things are happening at once, and they're pulling in the same direction: more stress, more deals, less bandwidth to sort through vendors carefully.
On one side, PwC's mid-year 2025 data put property values down 20% from peak, which means lenders and asset managers are carrying more distressed positions with staff counts that haven't grown to match. On the other side, the Mortgage Bankers Association logged $498 billion in total CRE lending in 2024, up 16% year-over-year, with Q4 2024 originations up 84% year-over-year and a 2025 forecast of $583 billion. So the same teams managing a wave of workouts are also underwriting a surge of new volume. Something has to give, and for a lot of firms, that something is due diligence time on the tools they buy to handle the load.
Meanwhile the vendor market is not exactly short on options. CRETI tracked roughly $16.7 billion deployed into CRE and infrastructure technology in 2025, a 67.9% jump year-over-year according to theaiconsultingnetwork.com. That's a lot of funded startups showing up at the same conferences, pitching the same credit committees, all promising to fix the exact bandwidth problem described above. Funded is not the same as proven, though, and a flush cap table says nothing about whether the underlying model can spread a T12 correctly. More distress, more volume, more vendors, all landing on desks at once: that combination is precisely what produces a rushed procurement decision, and a rushed decision is where the risk in this whole exercise actually lives.
Why generic horizontal AI tools fail CRE-specific workflows
Here's the thing about general-purpose AI: it's trained on broad slices of the internet and general business documents, not on the particular grammar of a rent roll or a loan agreement. CRE paperwork has its own vocabulary, its own structural quirks, and a long tail of edge cases that a model built for generic office work simply hasn't seen enough of.
What does that failure look like in practice? A model misclassifies a line item during financial spreading and produces a net operating income figure that's off, or an operating expense ratio that doesn't tie back to the source document. A lease has jurisdiction-specific clause language that a CRE-trained system would flag automatically, and a generic one just skims past. Or the tool generates a perfectly confident-sounding paragraph about covenant status without actually understanding how CRE covenants get structured or tested in the first place. None of these show up as an error message. They show up as a plausible-looking answer that happens to be wrong.
That plausibility is the real danger. The 2025 Keyway survey found investment committees already carry some skepticism toward AI-generated analysis, and a tool that produces confident, wrong CRE output doesn't ease that skepticism, it hardens it, often after the error has already done damage. Morgan Stanley has estimated that 37% of CRE tasks could be automated today, but that number says nothing about which tasks, by which tools, at what accuracy. A platform can sail through a pilot and still be unfit for production if the pilot never stress-tested it against the ugly, inconsistent document types that make up a real portfolio. Domain fit isn't a nice-to-have feature. It's a risk variable, and it belongs in the diligence framework the same way credit risk or interest rate risk would.
Output quality and accuracy standards a vendor must be able to demonstrate
Ask one question before anything else: can the vendor prove accuracy on documents that look like the buyer's actual portfolio, not on a curated demo deck built to make the product shine?
A real evaluation demands accuracy benchmarks on the document types a firm actually handles: rent rolls, T12 operating statements, loan agreements, draw requests. It demands error rate disclosure that goes past a headline accuracy number, down into false negative rates on covenant triggers and financial anomalies specifically, since those are the errors that cost money. And it demands source citation. Every output should trace back to the specific document passage or data field that produced it. An AI conclusion nobody can trace isn't analysis; it's a guess with good formatting.
Financial spreading makes a good stress test precisely because the correct answer is checkable. Hand the vendor a set of the firm's own financials, have them spread it, then compare the output line by line against what an analyst would produce by hand. Automating spreads can deliver meaningful efficiency gains, but only when the extraction underneath is accurate; speed without accuracy just means the errors compound faster and at greater scale. Watch for accuracy claims built on vendor-controlled test sets, figures that don't specify document type or complexity, and demo environments running on inputs that were cleaned up before anyone saw them. In a regulated lending environment, every AI output tied to risk or compliance needs to survive an examiner's questions. A vendor who can't explain how a number was derived cannot clear that bar, no matter how good the number looks.
How to evaluate covenant monitoring and risk detection capabilities specifically
Covenant monitoring earns its own line item in the diligence process because the stakes and the timing both work against periodic review. CREFC data from December 2025 put the effective CMBS delinquency rate around 8.75% once performing matured balloons are counted in, meaning lenders are sitting on more at-risk positions than a quarterly review cycle can realistically track. Risk doesn't wait for the calendar. A DSCR sliding from 1.31x to 1.28x to 1.26x across three quarters is telling a story well before it crosses whatever line counts as a breach, and a system that only checks thresholds misses the story entirely.
So what separates real monitoring from a glorified reporting dashboard? Continuous or near-continuous testing against live financial data, not a batch job that runs once a month. Trend detection that flags a deteriorating trajectory before the breach happens, not after. Portfolio-level exception reporting, since a lender holding hundreds of positions needs the system to surface what matters rather than requiring loan-by-loan review of everything. And records that hold up if an examiner asks to see them.
Portfolio monitoring and early warning already rank as the most prioritized generative AI use case among credit risk leaders, cited by nearly 60% of those surveyed, which means buyers should expect a vendor to show up with a mature, documented capability here rather than a roadmap slide. Some deployed systems have cut monthly review time from over 200 hours to under 30, with early alerts catching problems before they turned into write-downs; ask any vendor pitching this space for comparable numbers from actual client portfolios, not a projection built in a spreadsheet. Push on the specifics too: which covenant types does the system test, and what happens with the non-standard ones? How does the alert logic tell a genuine early warning apart from noise? And what happens when the source financials show up late, incomplete, or formatted in some way the system wasn't built to expect? A platform built specifically for CRE lending, with spreading, covenant testing, and portfolio dashboards built around the way lenders actually work, sits in a different category than a generic loan management system with a monitoring tab bolted onto the side.
Data security, privacy architecture, and what institutional-grade actually requires
Rent rolls, borrower financials, loan agreements, portfolio positions: these are among the most sensitive documents a firm produces, and any AI vendor touching them becomes part of that firm's data risk surface whether anyone thinks of it that way or not.
"Enterprise-grade security" needs to mean something specific, not just a phrase on a sales deck. Client data should never train a shared model or bleed across tenants; that's a baseline, not a premium feature. Encryption needs to cover data at rest and in transit, with key management the client can control or at least audit. Access controls and audit logs need to show who touched what, when, and what came out of it, which matters for internal governance and for any regulator who comes asking. And when a vendor claims SOC 2 Type II or ISO 27001, ask for the actual report. A self-attestation is not a certification.
One question deserves a direct answer, not a dodge: does the vendor use client documents to improve its own models? This happens more often than it gets disclosed, and it can create confidentiality problems or proprietary data exposure that nobody signed up for. CRE lenders also operate under examination by federal and state regulators, and AI tools feeding into credit decisions can fall under model risk management guidance, the kind that traces back to frameworks like SR 11-7 and whatever has followed it. A vendor needs to be able to speak to that directly, not wave it off as somebody else's problem. Request a third-party penetration test summary, a data processing addendum, and a clear statement of what happens to the data once the contract ends. A security failure in an AI platform isn't just an IT incident to route to the help desk; it's a reputational event and a possible regulatory violation, and no feature set is worth trading that away.
Whether the platform grounds AI outputs in the firm's own IP and templates
Every CRE firm has spent years building up its own underwriting templates, its own credit memo structure, its own asset-level history and deal logic. That accumulated knowledge is real competitive advantage, and a generic AI tool walks right past all of it.
Grounding in firm IP means something concrete. It means financial spreading maps to the firm's own chart of accounts instead of forcing a vendor-default schema that someone then has to manually reconcile back to the format the firm actually uses. It means covenant monitoring runs against the firm's real loan documents and its own internal thresholds, not a generic list of covenant types that only sort of matches. It means outputs speak in the firm's own risk language and escalation logic, so an AI-generated summary slots into the existing committee process instead of creating a second, parallel review nobody asked for.
The questions worth asking a vendor are direct: can the platform ingest the firm's specific templates and document formats starting on day one? How do firm-specific definitions and thresholds get maintained as the portfolio changes over time? And when the system hits a firm-specific edge case, is there a way to correct it, and does that correction stay inside the firm's own environment rather than feeding back into some shared model everyone else uses too? Skip this step and the alternative shows up fast: a platform that imposes its own schema forces analysts to translate every output back into the firm's framework by hand, which adds a manual step that erases the efficiency gain and opens a second door for errors to slip through. Agentic AI platforms that pull data from multiple sources and run multi-step workflows on their own are only as good as the institutional knowledge sitting underneath them. Without that firm-specific grounding, an autonomous output is generically fine at best and quietly wrong for this particular firm at worst.
Workflow integration and the pilot-to-scale failure pattern
Here's a pattern worth naming: a pilot finishes on schedule, hits every metric it was supposed to hit, and twelve months later the platform is still running on three assets or one region. No new integrations, no new teams using it, nothing. It just sits there, technically successful and practically stalled.
Why does this keep happening in CRE specifically? Pilots run on clean, hand-picked documents. Production runs on whatever actually lands in the inbox, which means inconsistent formatting, missing fields, and document structures the tool never saw during testing. Pilot teams also tend to be early adopters, the analysts already curious about new tools, while a full rollout means asking analysts with established habits and existing workflows to change how they work. And integration with the systems already in place, the loan management platform, the portfolio database, the reporting stack, often gets glossed over during the pilot phase and turns into the actual blocker once someone tries to scale it.
A few questions surface this risk before the contract gets signed. What are the vendor's documented integration points with the systems already running at the firm? How does the platform handle the document formats and input quality that fall outside whatever it was trained on? And what does implementation support actually look like once the initial rollout is done, not during the sales process, but six months in when something breaks? The right pilot doesn't just test the clean cases; it runs on the firm's own messy, non-standard documents, against production-readiness criteria decided in advance, not just proof-of-concept metrics chosen to make the demo look good. A platform built by people who've actually run CRE workflows should be able to name, from that experience, exactly where integration friction tends to show up and how they handle it. A vendor who can't get specific about that hasn't been through a production deployment at scale, and that's worth knowing before, not after.
Vendor credibility criteria that go beyond reference calls
Reference calls have a built-in limitation worth naming plainly: the vendor picks the references, the references are almost always success stories, and none of them are going to volunteer the failure modes a buyer actually needs to hear about.
So look past the call list for structural signals instead. Domain lineage matters: was this platform built by people who actually worked CRE transactions, or by technologists who took a generic AI architecture and mapped CRE terms onto it after the fact? That distinction shows up in how the product handles edge cases, ambiguous documents, and the specific logic of CRE risk, the parts that are hardest to fake and easiest to spot once someone knows to look. Client profile fit matters too: does the vendor's existing customer base actually resemble the buyer, in institution type, portfolio size, and workflow complexity, or is the buyer effectively signing up to be the design partner for a product that's never been tested at this scale before?
And transparency counts for more than most buyers give it credit for. A vendor who can say plainly what the platform doesn't do, where it needs a human in the loop, where the accuracy drops off, is telling a buyer something a polished pitch deck never will. That kind of honesty is rare enough in vendor conversations that its absence should probably count as a data point on its own. None of this replaces the work laid out in the sections above, the accuracy testing, the security review, the covenant monitoring specifics, the integration questions. It's the filter that decides which vendors are worth putting through that work in the first place.


