Est.

Evaluating AI Underwriting Vendors for CRE

Lenders need a framework for picking CRE-specific vendors, not generic AI platforms.

Features Editor · · 13 min read
Cover illustration for “Evaluating AI Underwriting Vendors for CRE”
CRE Tech Stack · September 1, 2026 · 13 min read · 2,814 words

Evaluating an AI underwriting vendor for commercial real estate is a different exercise from picking enterprise software for accounting or HR. The category is crowded with tools that call themselves "AI underwriting" while sharing almost nothing in terms of depth, and the criteria that actually separate a CRE-fit platform from a repositioned horizontal tool have little to do with the speed and accuracy claims every vendor puts on the front page of their website. This piece lays out that decision framework, focused less on ranking winners than on the questions a buyer needs to ask before signing anything.

Where the CRE market stands and why speed and accuracy now carry direct financial weight

Total U.S. commercial and multifamily mortgage borrowing hit $498 billion in 2024, up 16% from the year before, according to the Mortgage Bankers Association. The MBA's own forecast puts 2025 at $583 billion. Q4 2024 originations alone jumped 84% year-over-year, which means the volume of deals moving through underwriting desks spiked hard and fast, and any bottleneck in that pipeline now shows up as a real cost rather than a rounding error.

Layer onto that the margin picture. CBRE reported commercial mortgage spreads compressed 49 basis points year-over-year in Q4 2024, landing at 184 bps. Thinner spreads mean the money isn't there to absorb inefficiency the way it once was. A shop that takes three extra days to spread financials on a deal is leaving basis points on the table in a market where basis points are the whole game. When a vendor pitch leads with "we make underwriting faster," the useful response is a question about whether faster actually shows up on the P&L given today's spread environment. Workflow integration and decision speed belong in the evaluation criteria alongside feature lists, factored into a demo from the start.

The adoption gap that makes vendor selection consequential right now

Here's a stat worth sitting with: JLL's 2025 Global Real Estate Technology Survey found that 92% of CRE firms have piloted AI in some form, but only 5% report hitting all their AI goals. That gap between "we tried it" and "it worked" has as much to do with whether what got piloted was built for what CRE underwriting actually demands as with whether the technology functions at all.

Adoption is real, just lopsided. The same survey period saw institutional investor use of AI for market analysis climb to 61% in 2025, up from 22% in 2023. That's genuine movement. But market analysis is a lower-stakes task than underwriting a loan that a credit committee has to sign off on, and a separate 2025 survey from Keyway found investment committees distrust AI-generated underwriting analysis at meaningful rates, with only a minority expressing real confidence in using it for that specific purpose.

Is that distrust irrational? Not particularly. A credit committee's job is to ask "how do you know that," and if the honest answer is "the model said so," the committee has every reason to push back. That points to a product design problem as much as a trust problem, and it shows up again later in this piece when the conversation turns to source citations. For now, the practical consequence is this: firms still running vendor bake-offs are competing against early movers who already have AI running across the full deal lifecycle, and that window to catch up is closing.

Why CRE-native depth is the first criterion, not an afterthought

Picture three categories of tool sitting side by side. One is an AI-native CRE platform built to extract data from whatever document format shows up and populate a model automatically. Another is legacy underwriting software that still wants an analyst typing numbers into fields by hand. The third is a horizontal AI tool, built for general document processing, that got a CRE label slapped on it sometime after a funding round.

The third category is where things go wrong in specific, avoidable ways. A generic AI model doesn't know that a rent roll's "other income" line usually excludes utility reimbursements, or that a DSCR schedule structured for a hospital lease reads nothing like one for a warehouse. Feed a horizontal tool a multifamily rent roll and it might classify a concession as revenue, misread a lease step-up as a one-time charge, or miscalculate NOI because it doesn't know which operating expenses belong below the line for that property type. These are the ordinary texture of CRE documents rather than edge cases, and a model without CRE-specific training treats ordinary texture as noise.

What does CRE-native depth actually look like when it's working? A handful of things, concretely. The platform should understand property-type-specific financials for multifamily, office, retail, and industrial without a custom setup process. It should read rent rolls, DSCR schedules, appraisals, operating statements, lease abstracts, and environmental reports as native formats, without needing pre-processing. It should populate the firm's own underwriting templates and Excel models, keyed to that firm's conventions rather than a generic model the vendor built for nobody in particular. And every extracted figure needs a citation back to its source document, so the credit committee can check the work instead of taking it on faith.

There's a useful diagnostic question buried in all this: who designed the underwriting logic, and have they actually closed CRE deals? A platform built by people who spent years structuring loans encodes assumptions about how deals actually move that a team of computer scientists, however talented, simply hasn't lived through. Underwriting logic built without transaction scars tends to miss the scars that matter.

What documented efficiency gains actually tell you — and what they don't

The efficiency numbers floating around the industry are genuinely striking. Time-to-decision reductions in the 50 to 75% range have been reported for banks deploying AI underwriting on commercial loans. CBRE's 2025 Technology in Real Estate report found AI-adopting firms completing deal analysis 60% faster than peers still running manual workflows. Vendor-reported figures for underwriting agents show cycle times falling around 41%, throughput reaching roughly three times per analyst without adding headcount, and financial spreading time down about 36%. On the cost side, published lender data points to savings up to 40% in loan processing, around $1,500 per loan (a 14% reduction), and production cycles shortened by about five days.

Those numbers are real, and they're also, almost without exception, vendor-generated or vendor-commissioned. That doesn't make them false. It does mean the sample size, the deal types tested, and the baseline workflow being compared against can vary enormously from one study to the next, and a vendor has every incentive to publish the comparison that flatters them most.

The useful move is to interrogate the numbers rather than accept or dismiss them outright. What was the baseline: a manual spreadsheet process, legacy underwriting software, or a competing AI tool? What property types and document formats made up the test set, and do they resemble what actually crosses your desk? Were the results pulled from a curated pilot with clean documents, or from production conditions with the scanned PDFs and inconsistent formatting that real deal flow produces? And can the vendor put you in touch with a reference client whose deal mix looks something like yours, rather than their best-case customer?

The question worth asking in a vendor meeting isn't "how fast is your platform." It's "how fast on deals like ours, with documents like ours, at the volume we actually run." That specificity is what separates due diligence from marketing.

Financial spreading as a concrete test of what a platform can actually do

If there's one place to run a real test rather than take a claim at face value, it's spreading. Every vendor says their platform automates it. Far fewer can show accurate, consistent, auditable output when the input documents aren't pristine, and spreading happens to be the one part of underwriting where quality is genuinely measurable, not just asserted.

The manual version of this problem is familiar to anyone who has done the work: an analyst pulling figures by hand from operating statements, rent rolls, and borrower financials into a model, introducing small inconsistencies along the way and burning hours that would be better spent actually analyzing the deal. A platform that's doing this well handles tax returns, P&Ls, balance sheets, audits, and 10-Ks without needing them pre-formatted first. It categorizes line items against the firm's own chart of accounts, keyed to that client's setup rather than a generic template. It calculates DSCR, debt yield, and related ratios while showing the data path underneath, and it flags inconsistencies (rent roll income disagreeing with P&L income, say) instead of quietly picking one and moving on. A platform that handles tax returns, audits, company-prepared statements, and 10-Ks and 10-Qs without pre-formatting is a reasonable benchmark for what document-type coverage should look like when comparing other vendors.

Here's the actual test, and it costs nothing but a little time: take a real, anonymized document package from a recent deal, mixed formats and multiple years included, and run it through the vendor's platform. Then compare that output against what an analyst produced by hand. The gaps that show up will tell a buyer more in twenty minutes than any scripted demo will in an hour.

One distinction matters more than it might seem at first: a platform that populates its own generic model is doing a different job from one that populates the firm's own templates using the firm's own definitions. The second version produces output someone can actually use without redoing half the work.

Covenant monitoring and portfolio surveillance as ongoing operational criteria, not just deal-time features

Most vendor evaluations obsess over deal-time speed and quietly skip the question of what happens after closing, which is a mistake, because the more durable value tends to show up in surveillance as much as origination.

The legacy approach to covenant tracking is a spreadsheet somebody updates once a quarter, which means breaches tend to surface only after they've already happened, not while there's still time to do something about them. That timing gap matters more than it sounds. A technical default, a covenant breach even with every payment current, can trigger a cash sweep, default interest, or acceleration of the loan. Catching a DSCR trend sliding downward two months ahead of the test date gives a lender room to negotiate; catching it on the test date itself does not, and the difference between those two outcomes is entirely a function of whether the monitoring system watches continuously or checks in occasionally.

Built Technologies offers a useful reference point for what continuous monitoring looks like at real scale: more than 300 lenders managing over $317 billion in construction and CRE loans, with reported capacity gains of 2 to 5 times and meaningfully faster draw processing. That's the kind of number that suggests portfolio-wide surveillance is already running in production somewhere, grounded in real deployments rather than a hypothetical capability.

The questions worth asking a vendor here: does the platform track DSCR, LTV, debt yield, and reporting covenants continuously, or does it run on a scheduled batch that might miss something between check-ins? Does it surface a deteriorating trend, or only a confirmed breach after the fact? Can it flag portfolio-level patterns, loans approaching maturity, refi windows opening up, concentration risk building in one submarket, without someone manually reviewing every deal? And can the alert thresholds be configured to match the firm's own covenant definitions, tuned to that firm's risk appetite rather than a generic default?

Data grounding, source citations, and why "the AI said so" is not an acceptable audit trail

Go back to that JLL trust gap for a moment. It reflects a product gap as much as a change management issue that better training might address. If a platform can't show exactly where a number came from, no investment committee should be expected to rely on it, and no amount of internal advocacy changes that math.

Source-cited output means something specific: every extracted figure links back to a page and line in the original document, an analyst can check the work without rereading the entire file, and the resulting credit memo is auditable by a committee member who wasn't even in the room when the model ran. That last part matters more than it sounds, because credit committees operate on the ability to verify.

There's a related distinction worth separating out clearly: a platform drawing on its own generic knowledge base is answering a different question than one grounded in a firm's actual documents, templates, past deals, and internal data. The first produces analysis. The second produces analysis that matches how the firm actually underwrites, which carries a different kind of reliability altogether.

Worth asking directly: is the output grounded in the documents uploaded, or is the model pulling in outside training data that can't be inspected? Can the platform take on the firm's own templates and definitions as its reference point, tuned to how that firm already works? And when the model hits something it's genuinely uncertain about, does it flag that uncertainty, or does it just fill the field and move on, leaving the uncertainty invisible until someone catches it downstream? One reference point worth noting: Hypha, a CRE-focused AI underwriting platform, grounds its output in each firm's own templates and uploaded documents rather than a generic knowledge base. Henry.ai builds a living database from a firm's own deals and runs underwriting through the firm's existing Excel models via a two-way add-in, keeping the firm's own institutional knowledge central to the process.

Security and data privacy requirements that CRE institutions should not negotiate away

CRE underwriting touches borrower tax returns, personal financial statements, proprietary deal terms, and portfolio data that a lender has every reason to keep locked down. That's a materially higher sensitivity level than most horizontal SaaS tools were ever designed to handle, and it's worth treating the security conversation as a priority from the start of procurement rather than a box to check late in the process.

A short list of things to require, not just ask about: SOC 2 Type II certification specifically, because Type II means the controls have actually been tested over a period of time rather than assessed on paper once (a self-attestation or Type I report falls short of that bar). Data residency and tenant isolation, so a firm's data isn't sitting in the same bucket as a competitor's. Clarity on model training boundaries, because if a vendor trains or fine-tunes its models on client documents, a firm's deal data may be quietly shaping outputs a competitor later sees. And role-based access controls, so not everyone inside the firm sees everything inside the platform.

SOC 2 Type II certification gives buyers a useful baseline for what to expect from serious vendors in this space, and confirming that certification should be a standard step in any procurement process. The single most revealing question to ask in a vendor call is blunt: does our data ever leave our tenant environment to improve your model? The answer determines whether sensitive borrower information stays protected or ends up, indirectly, training a shared model that other clients benefit from. For lenders under bank examination or CMBS reporting obligations, that answer carries regulatory weight too, and a vendor that can't produce a clean audit is a liability sitting in the vendor stack.

Integration, implementation depth, and what "easy to deploy" usually means in practice

Every demo runs on a curated set of clean documents chosen specifically because they make the platform look good. Production looks different: mixed formats, incomplete files, exports from a legacy loan origination system that nobody fully understands anymore, and hand-offs between teams that a sales demo will never show a prospective buyer.

The real integration question is whether the platform connects to systems already running in the shop, things like Yardi, MRI, Argus, or an internal loan origination system, or whether it creates a second, parallel workflow that analysts now have to maintain alongside everything else. Northspyre's integrations with Yardi, MRI, and Sage are a reasonable reference point for what deep accounting-level integration actually looks like, going beyond a surface-level API connection that technically works but doesn't move real data.

And then there's the document question that decides whether any of this holds up outside a sales pitch: can the platform actually ingest what borrowers and brokers really send, the scanned PDFs, the handwritten amendments scrawled in a margin, the multi-entity financial packages that never come in a clean single file? That gap separates a platform that performs well in a demo from one that performs well on a Tuesday afternoon three months into a live deployment, and it's the last, most practical filter in a framework that, at every stage, comes back to the same basic instinct: ask for proof, on real documents, at real volume, before the contract gets signed.

Sources

  1. smartcapitalcenter.com
  2. smartcapitalcenter.com
  3. thefractionalanalyst.com
  4. getbuilt.com
Filed underCRE Tech Stack

More in CRE Tech Stack