Est.
AI in CRELong read

AI Risks and Limitations in CRE Underwriting

Features Editor · · 12 min read
Cover illustration for “AI Risks and Limitations in CRE Underwriting”
AI in CRE · August 6, 2026 · 12 min read · 2,678 words

There is something almost poetic about an industry that prices long-duration illiquid risk for a living being among the last to govern the tools it uses to do it. Commercial real estate underwriting has historically been part science, part séance: you are discounting cash flows that do not yet exist, on assets that do not trade frequently, in markets that change faster than any model updates. And yet here we are, deploying generative AI into that workflow and acting surprised when the outputs require supervision.

That surprise is the problem worth examining. Not because AI has no place in CRE underwriting; it clearly does. But because the gap between what AI is good at and what underwriting actually requires is wide enough that conflating the two creates a different kind of risk than the one you thought you were solving. Total commercial mortgage origination is projected to exceed hundreds of billions of dollars in 2026, per the Mortgage Bankers Association's 2025 Commercial Real Estate Finance Forecast. That volume is pushing institutions toward AI faster than their governance frameworks can follow. The Altus Group's 2025 CRE Innovation Report documented a significant jump in production AI adoption over two years, yet NAIOP's 2024 CRE Technology Adoption Survey found that a majority of firms piloting AI had not achieved their program objectives. The gap lives in underwriting, where the decisions are real and the errors get funded.

The argument this piece wants to make, and wants you to push back on, is this: some AI risks in underwriting are structural, meaning they cannot be engineered away with better prompting, richer data, or a more sophisticated model. Others are solvable, meaning they are process and design problems that respond to specific choices. Conflating the two categories leads either to over-reliance or to reflexive rejection. Neither is a strategy.

Venn diagram: AI in CRE Underwriting: Structural vs. Solvable Risks. Compares Structural AI Risks and Solvable AI Risks; overlap: Shared Risks.

What AI Is Actually Doing in a CRE Underwriting Workflow Today

To understand where AI fails, you first have to understand where it is actually deployed, because the risk profile of spreading a tax return is not the same as the risk profile of drafting an investment committee memo.

The core problem AI addresses in underwriting is not credit analysis; it is document processing. Analysts have historically spent the majority of their working hours on extraction: spreading tax returns, reconciling rent rolls, chasing down operating statements in three different formats from three different property managers. The credit analysis, the part that requires judgment, comes after. AI accelerates the before.

What AI-powered spreading actually does is ingest structured and unstructured inputs, PDFs, scanned returns, Excel interim statements, classify line items, align periods, reconcile totals, and flag anomalies, outputting a model-ready dataset rather than a raw document pile. That is useful. Beyond spreading, AI is being used for rent-roll extraction, pro forma generation, risk scoring, covenant extraction, and, in more advanced deployments, drafting IC memos. Agentic frameworks that coordinate sequences of tasks across a deal lifecycle without manual handoffs are beginning to appear in production, though they remain early.

What AI is not doing in most institutional deployments: making final credit recommendations, replacing analyst judgment on sponsor quality or market trajectory, or signing off on exceptions. That task map matters enormously for risk assessment, because the failure modes of AI doing extraction are different in kind from the failure modes of AI generating a risk score, and both differ from AI drafting a memo that a fatigued analyst reads once and forwards to committee.

The question is not whether AI belongs in the workflow. It clearly belongs in parts of it. The question is which parts, and what happens at the boundaries.

Hallucination in a Context Where Every Number Is Load-Bearing

Most coverage of AI hallucination treats it as a novelty problem, an awkward anecdote about a chatbot inventing a court case. In underwriting it is a balance-sheet problem.

Hallucination is not a glitch. It is a property of how large language models generate output: they produce plausible text, not verified truth. The model cannot distinguish between a figure it extracted from a source document and one it synthesized from context because both feel, to the model, like the same act of generation. In most domains this produces an embarrassing citation. In underwriting it can produce a DSCR that never existed, a covenant threshold that was fabricated, or an NOI figure silently pulled from the wrong year's financials.

Why does this happen at a structural level? Generic models are not trained on a firm's specific templates, rent roll formats, or document conventions. They have no internal mechanism to flag uncertainty about whether a figure appears in the source or was inferred from plausible surroundings. Without source citation — a traceable link from every output figure to a specific location in a specific document — there is no way to audit the output before it enters the model.

The solvable dimension here is real: hallucination risk drops sharply when AI is constrained to extract rather than generate, pulling figures from source documents with citations rather than reasoning from context. When outputs are designed to be auditable against the original, the analyst can verify rather than accept. That is an engineering choice, not a fundamental limitation.

But the governance implication cuts in both directions. Firms that deploy AI with output traceability have meaningfully reduced a specific risk. Firms that deploy AI without it have not reduced underwriting risk; they have added a new upstream failure mode to a process they believed they were improving. The spreadsheet error you could trace was safer than the AI output you could not, because at least the former was auditable.

Training-Data Gaps and What Historical Data Cannot Capture About CRE Assets

Even well-constrained, well-cited AI runs into a second structural problem: the training data that underlies CRE models is thin in the places that matter most.

Consider what "thin" actually means in this context. Specialty assets, secondary and tertiary markets, and distressed situations with complex capital structures all have limited transaction histories. Those are often precisely the deals where lenders most need reliable pattern recognition — the unusual situations where an analyst's judgment is least reliable and where a well-trained model would theoretically add the most value. The data is thin exactly where the analytical need is greatest.

The temporal problem is a related but distinct failure mode. Trailing-twelve-month financials and historical comps are the inputs most AI models are trained on. They say nothing about a lease expiring in 18 months, a tenant on a lender's internal watch list, or a submarket quietly absorbing new supply that has not yet shown up in reported vacancy. AI-powered automated valuation models perform reasonably well on stabilized commercial assets; Green Street's 2025 AVM Benchmark Study found median error for institutional-grade models on that specific asset class around ±4.8%. That figure degrades materially for specialty properties, distressed assets, and markets with thin comps. The model is most accurate where analytical support is least needed.

It is also worth considering what the historical data encodes beyond simple measurement error. Training sets carry historical pricing patterns, and those patterns reflect past market behavior, not forward conditions. Certain submarkets, property types, or borrower profiles are systematically underweighted or overweighted because the model learned from a market that no longer exists. That is not a data pipeline problem you can solve by adding more comps. The fundamental heterogeneity of CRE assets means thin-data conditions are a persistent feature of the asset class, not a temporary gap to be closed.

What this means structurally is that AI performs best where data is dense and assets are homogeneous. Those are often not the deals where analytical support is most needed.

The Judgment Gap That No Model Currently Closes

One might argue that as models improve, as training data grows richer and more current, the judgment gap narrows toward zero. That is worth taking seriously, and then questioning for a specific reason.

CRE underwriting is not purely a data problem. The final credit decision integrates factors that are rarely written down anywhere. What experienced underwriters are actually doing when they read a deal: assessing a sponsor's execution track record, and not just the resume, but what that resume means in this specific market cycle. Reading local market conditions that do not appear in any data feed: a city council vote, an anchor tenant's planned exit, a neighborhood trajectory visible on foot but not yet in the comps. Weighing the qualitative difference between a signed LOI and a signed lease, which is a judgment call that depends on knowing the tenant, the market, and the broker.

JLL's Becci Curry, as reported by Bisnow, noted that senior practitioners want a human accountable party — someone, not just an AI, to be responsible if something goes wrong. That observation is sometimes dismissed as institutional conservatism. It is actually a description of how accountability structures function. Credit decisions carry personal professional liability. AI does not.

The failure mode to watch is not that AI makes a bad recommendation. It is that analysts treat a model output as a conclusive answer and stop doing the interpretive work that catches what the model cannot see. Over-reliance is insidious precisely because it looks like efficiency. The analyst finishes faster. The deal moves faster. The error surfaces later.

This is not a limitation to be engineered away. It reflects the fact that CRE assets are one-of-a-kind properties in specific markets with specific sponsors, and the relevant judgment context is not reducible to structured data. AI is best positioned as a decision-support layer, not a decision layer. The goal is freeing analysts to spend more time on the judgment work, not replacing it.

The Governance Gap That Turns Manageable Risks Into Institutional Ones

That raises an important question: if the structural risks are understood and the solvable risks are tractable, why are so many deployments still producing poor outcomes?

The answer is governance, or rather its absence. Most CRE lenders have access to AI tools before they have the frameworks to deploy them responsibly. Data lineage, approval workflows, output traceability, model validation: these are not glamorous problems and they do not feature in vendor demos. Per Deloitte's 2025 Model Risk Management Survey, more than half of banks cite transparency and explainability as a primary hurdle to AI deployment — not the tools themselves but the inability to demonstrate how a conclusion was reached. Regulators examining an underwriting decision want a paper trail. AI that cannot produce one creates exposure that did not previously exist.

Forrester's 2025 AI Governance Forecast projected that ungoverned generative AI in commercial applications will cost well into the billions; the cost is not the technology but the absence of controls around it. The deployment failure pattern is recognizable to anyone who has watched a technology rollout from the inside: a firm distributes AI tools broadly, asks teams to find efficiencies, and discovers a year later that most of the effort went into building controls around the tool rather than using it. The governance work consumed the operational gain. Per Dealpath's 2024 survey of buy-side investors, lack of infrastructure accounts for the majority of failed AI deployments — not the AI itself.

What adequate governance looks like in underwriting specifically: every AI output is traceable to a source document with citation; model assumptions are auditable, not black-box; human sign-off is a structural step, not an optional review; and the firm's own underwriting templates and standards govern what the AI produces, not a generic model's defaults. That last point is underappreciated. A model calibrated to the firm's credit culture, its definition of stabilized occupancy, its treatment of management fee add-backs, produces outputs that the firm's analysts can interrogate. A generic model produces outputs that look authoritative and are substantially harder to challenge.

The regulatory pressure is real. Citizens Bank's 2025 survey found that a large majority of respondents agreed that AI requires significant effort to identify legal and appropriate use cases, and regulators are watching underwriting AI specifically. The question is not whether oversight is coming; it is whether firms will have their governance structures in place when it arrives.

Where AI Risk Compounds After Closing: Covenant Monitoring as the Overlooked Exposure

Most AI investment in CRE focuses on the origination phase. That is where the demos are, where the efficiency gains are most legible, and where vendors compete. The monitoring phase is where risk actually accumulates over the life of a loan.

Per PwC's 2026 survey of credit portfolio managers, more than half prioritize AI in underwriting, but only a small fraction treat AI-enabled portfolio monitoring as a current priority. That gap is significant because commercial real estate loans rarely fail because a borrower misses a payment. They fail because a borrower quietly breaches a covenant — a DSCR floor, an LTV ceiling, a financial reporting deadline — months before a lender detects it. A mid-sized lender with hundreds of active commercial loans is tracking thousands of covenant thresholds each quarter, and a large majority of banks still manage that with spreadsheets and manual processes.

AI in monitoring addresses a different set of limitations than AI in underwriting. Extraction risk is lower; covenants are defined terms in executed documents. The risk is in mapping those terms to live borrower financials continuously, not just at origination. Early warning systems require continuous data ingestion and structured exception flagging, not a one-time spreading event.

The market structure makes this more urgent. Covenant-lite structures have grown substantially in private credit, EBITDA add-back caps have eroded, and payment-in-kind structures convert cash signals into silence. The remaining covenants matter more precisely because there are fewer of them. A breach in a lean covenant package is a serious signal. Missing it is a serious failure.

The firms that will absorb the most avoidable loss are those that invested in AI at origination and left monitoring manual. They accelerated into the loan and slowed down on the surveillance.

How Firms That Deploy AI Well Are Drawing the Line Between Tool and Judgment

The structural risks define where AI should function as a support layer, not because the technology will invariably be inadequate, but because the accountability structure of lending requires a human decision-maker. That is not a temporary condition pending a model upgrade. It is a feature of how credit responsibility is allocated.

The solvable risks respond to specific design choices. Hallucination without traceability is an engineering problem. Ungoverned deployment is a process problem. Governance theater — the appearance of oversight without the substance — is a leadership problem. All three are tractable.

The design choices that separate well-deployed AI from poorly-deployed AI in underwriting share a common logic: AI extracts and cites, humans interpret and decide. Outputs are traceable to source documents so that analysts can verify rather than accept. The AI is constrained by the firm's own underwriting templates and standards. Human sign-off is a structural checkpoint. Portfolio surveillance is continuous, not triggered only by origination events.

Firms that manage this line well are moving from document receipt to decision-ready analysis in a fraction of the time previously required, without introducing the new exposure that comes from ungoverned outputs entering credit decisions. The efficiency gain is real. The discipline required to capture it without creating new risk is also real, and it is less often discussed.

But what of the firms that skip the discipline? They have replaced manual inefficiency with AI output and called it risk management. The inefficiency is gone. The accountability structure was never rebuilt around the new process. That is not an improvement. It is a substitution of one failure mode for a faster one.

The question for any CRE institution deploying AI in underwriting is not "does it work?" Almost everything works in a demo. The question is: do we know exactly where it works, where it does not, and who is accountable when the line gets crossed? That question, answered honestly, is what separates firms deploying AI as a genuine improvement from those deploying it as a sophisticated way to feel current while the exposure compounds quietly downstream.

Sources

  1. bisnow.com
  2. blooma.ai
  3. aiforcrecollective.com
Filed underAI in CRE

More in AI in CRE