Est.

Human in the Loop Workflows for CRE AI

AI pilots are everywhere in CRE, but most stall because no one defined which decisions stay human.

Senior Writer · · 14 min read
Cover illustration for “Human in the Loop Workflows for CRE AI”
AI in Commercial Real Estate · July 29, 2026 · 14 min read · 3,245 words

There is something almost comic about the state of AI in commercial real estate right now. According to Keyway's 2025 State of AI Adoption survey, 92% of CRE teams have either started piloting AI or plan to do so this year, a near-total reversal from under 5% just two years prior. Yet per NAIOP research, only 5% of investors have actually achieved their program objectives. The industry has, collectively, sprinted to the starting line and stopped. It's like buying a gym membership in January and calling yourself an athlete by February.

The gap isn't a technology problem. The technology works. The gap is a model governance problem: most organizations have no structured answer to the question of which decisions AI makes, which it supports, and which stay entirely human. In an industry where a single commercial loan decision can move tens of millions of dollars, that ambiguity isn't a philosophical inconvenience. It's operational risk.

The Mortgage Bankers Association put 2024 U.S. commercial and multifamily mortgage lending volume at $498 billion, up 16% from 2023. Volume is rising. Velocity is rising. The decisions that AI is being asked to assist with are consequential, irreversible, and scrutinized by examiners. Even among well-capitalized institutions, McKinsey's July 2025 survey found that only 12% of North American banks had deployed any generative AI use cases at all. The chasm between pilot and production is vast, and the reason most organizations are stuck in it is that they haven't answered the governance question.

This article isn't a debate about whether AI belongs in CRE. That debate is over. The question now is what makes CRE AI deployments actually work, and the answer, consistently, is human-in-the-loop design.

What Human-in-the-Loop Means in Practice, and What It Doesn't

Let's be precise about terminology, because "human-in-the-loop" has become one of those phrases that means whatever the speaker needs it to mean. In its most useless interpretation, it means a human reviews everything the AI touches. That definition makes the technology pointless: you've added a tool that generates outputs and then assigned a person to re-examine every output before it moves anywhere.

The operative framework is a tiered decision structure. At the base tier sit fully automatable tasks: document parsing, routine data extraction, normalization, daily monitoring checks. These require no human approval because the consequence of a single error is low and the volume of items makes individual review impractical anyway. The middle tier covers outputs with material consequence: a flagged covenant breach, a spreading anomaly, a risk score embedded in a credit memo. These are AI-surfaced and human-reviewed, precisely because the downstream decision carries weight. The top tier is human-owned, full stop: final lending decisions, sponsor assessments, workout negotiations. Accountability that cannot be delegated shouldn't be.

The emergence of agentic AI raises the stakes on this design. When AI moves from answering questions to executing multi-step workflows autonomously, a shift the industry calls agentic AI, the decision layer doesn't become less important; it becomes more important. The system is now doing things, not just saying things. Where it hands off to a human, and what information it provides at that handoff, determines whether the human is exercising genuine judgment or rubber-stamping a fait accompli.

The governance framework that resolves this is explicit documentation: which tier each workflow occupies, with full traceability built into the architecture. Every input, output, and human approval, logged. That's not bureaucratic overhead. That's what makes a CRE AI deployment auditable, defensible to examiners, and trustworthy to counterparties and boards.

One more distinction worth drawing: oversight as bottleneck versus oversight as accelerator. The difference isn't whether approval paths exist. It's how they're designed. Approval paths calibrated to risk level compress cycle times without cutting corners. Approval paths designed as universal checkpoints just recreate the manual process.

The Opendoor Case and What Happens When the Loop Is Removed

Opendoor's early model included a sensible human verification step: local agents checked the AI's home price predictions before offers went out. Then scale pressure arrived, as it always does, and human verification was minimized. Offer-making was automated. The AI's predictions were taken at face value in a volatile market, and the result was systematic overpricing.

The instructive part of this story isn't that AI produced bad predictions. It's that the governance architecture was removed before the system had demonstrated sufficient accuracy to operate without it, and then the error propagated at scale before anyone caught it. The AI did exactly what it was designed to do. The design removed the correction mechanism.

Why does this apply with particular force to CRE lending? Because commercial real estate is less liquid, less comparable, and more structurally complex than residential. There are no truly equivalent transactions for a mixed-use, ground-leased asset in a tertiary market with a preferred equity stack. The AI's training data, however extensive, will always be operating in territory where local context, deal structure, and sponsor quality matter enormously and are not fully legible from documents alone. The margin for unverified AI output isn't wider in CRE than in residential. It's narrower.

The Opendoor case is the cautionary version of the story. The rest of this article is about what the constructive version looks like.

How Covenant Monitoring Illustrates Why Early AI Involvement and Late Human Judgment Is the Right Sequence

Consider the arithmetic of covenant monitoring at a mid-sized commercial lender. Five hundred active commercial loans, three to five covenants each: that's somewhere between 1,500 and 2,500 covenant thresholds to track every quarter. Approximately 70% of banks still manage this with spreadsheets and manual processes. The practical consequence is that covenant monitoring happens periodically, not continuously, and the loans that get reviewed carefully are often a function of which analyst happened to have bandwidth that quarter.

But here's the quieter problem. CRE loans rarely fail because a borrower missed a payment. They fail because covenant breaches accumulate undetected: DSCR slipping below 1.25x, LTV creeping above 75%, financial statements arriving 50 days after the required 45-day window. These are technical defaults that live in the loan agreements, not in the payment history. They're the early signals that, caught in time, give the lender options. Missed, they surface only after the situation has deteriorated to the point where options are limited and losses are probable.

AI changes this workflow materially, functioning as a continuous early warning system. Natural language processing extracts covenant terms from loan agreements; those terms are mapped to live borrower financials; every threshold is tested continuously, not quarterly; and negative trends are flagged before they become breaches.

But notice where the human judgment enters. The AI detects and flags early. The human reviews the flag and does the thing the AI cannot do: assess context. Is a DSCR (debt service coverage ratio) at 1.24x a one-quarter anomaly driven by a large one-time capital expense, or is it structural deterioration in the property's operating performance? That distinction determines the response: workout, waiver, or watch. The human makes that call. Think of the AI as a smoke detector — it tells you something is burning, but it takes a person to decide whether to call the fire department or open a window.

Consider a multifamily loan at $24 million with a 1.25x DSCR covenant. AI flags a declining trend at 1.24x two months before the test date. The borrower, now with time to act, accelerates three lease renewals and defers a capital project. DSCR recovers. Without continuous monitoring, that breach surfaces only when the lender activates a cash sweep, by which point the relationship, the property performance, and the lender's options have all deteriorated. One documented case involved a $200 million CRE lender that reduced monthly covenant review hours from over 200 to under 30; early alerts prevented two potential defaults, saving an estimated $1.2 million in write-downs.

That raises an important question: where does human sign-off become non-negotiable rather than merely advisable? The answer is examiner-readiness. The April 2026 interagency guidance requires an end-to-end audit trail: source document, page, formula, calculated value, and the identity of the human who approved the disposition. A system that flags and auto-resolves without a human decision in the chain doesn't satisfy that requirement. Human sign-off at the disposition stage isn't a concession to AI's limitations; it's a regulatory necessity.

One technical note that practitioners learn quickly: parsing an entire complex loan agreement in a single AI pass invites errors. Best practice breaks the analysis into focused phases, each verified before the next proceeds. The human reviewer is positioned to catch what a single-pass model confabulates, a failure mode practitioners increasingly call hallucination.

Financial Spreading and Underwriting: Where AI Handles the Grind and Humans Own the Judgment

Ask any commercial real estate underwriter where their time actually goes, and the answer is dispiriting: scrubbing PDFs. Operating memoranda, trailing twelve-month statements, rent rolls, tax returns, each arriving as an unstructured document that must be manually reconciled, keyed into a model, and cross-checked against the prior period. This data ingestion phase is where the one-to-four week CRE deal cycle mostly lives. The actual analysis, the part that requires professional judgment, is a fraction of the time.

AI financial spreading targets this bottleneck directly. Relevant figures are extracted from unstructured documents while preserving footnote context; metrics are calculated and benchmarked against peer and industry data; anomalies and internal inconsistencies are flagged automatically. For lease abstraction specifically, a 40-page lease that takes an analyst 45 to 90 minutes to abstract can be processed in two to three minutes. A 50-lease portfolio that would consume 35 to 70 hours of analyst time becomes a morning's work.

Banks using AI-assisted underwriting have reported 50 to 75% reductions in time-to-decision for commercial loans, with some lenders underwriting three to four times more deals with the same team. A McKinsey and IACPM December 2024 analysis of multiagent AI systems in credit workflows found analyst productivity gains of 20 to 60% and roughly 30% faster credit decision turnaround, with the analyst's role shifting from manual drafting to strategic oversight and exception handling.

But what does the hybrid model actually look like in practice, rather than in the vendor pitch? Excel isn't going anywhere. AI handles unstructured data processing, document analysis, and pattern recognition; Excel retains its place for structured financial modeling, custom formula logic, and auditable calculation chains. CRE investors using integrated AI and Excel workflows close deals meaningfully faster than those relying on either tool exclusively. The tools are not competitive; they're complementary.

It is also worth considering where human judgment remains genuinely irreplaceable, because the list is specific. Evaluating sponsor quality and track record: no document set fully captures whether a borrower will perform under stress. Assessing market-level context: AI reads the rent roll, but it doesn't know that the anchor tenant two blocks away just vacated or that the submarket's absorption rate turned negative last quarter. Scenario stress-testing interpretation: AI runs the models efficiently; humans decide which scenarios are plausible, which are extreme, and what the results imply for credit appetite. Final credit decision and accountability: this one is both irreplaceable and non-delegable.

One more data point worth sitting with: AI underwriting has demonstrated 15 to 30% higher accuracy in predicting defaults compared to manual-only processes. This matters for HITL design because the risk calculus is symmetric. Removing the human layer from the final decision carries governance risk. Removing AI from the analytical process carries accuracy risk. Good HITL design accounts for both.

Portfolio-Level Intelligence: The HITL Opportunity That Most Firms Haven't Reached Yet

Deal-level AI is where most firms have focused their early deployments, and reasonably so: the use cases are concrete, the time savings are measurable, and the governance questions are relatively tractable. But the more structurally significant opportunity is portfolio-level intelligence, and most institutions haven't gotten there yet.

The manual portfolio review problem is subtle but consequential. When every loan is reviewed on a rotating or periodic basis, the opportunities and risks you discover are a function of when review happened to fall, not of where risk actually is. A loan that was reviewed three months ago has deteriorated materially in the interim; a loan reviewed last week looks fine on paper but is second in the queue for a sponsor who's managing liquidity across six assets simultaneously.

The maturity wall concentrates this problem. Substantial volumes of CRE and multifamily mortgages were set to mature in 2025, creating a portfolio-level event that no manual process can fully monitor: refinancing windows, maturity triggers, and distress signals across hundreds of loans simultaneously. The lenders with continuous, AI-powered portfolio surveillance had options; the lenders relying on periodic manual review were reactive.

What portfolio-level AI actually does: it continuously surfaces loans approaching covenant thresholds across the entire book; flags maturity dates and refinancing windows before they become urgent; identifies clusters of exposure, geographic concentrations, sponsor-level exposure, asset-class concentration risks, that don't appear when reviewing loans individually; automates document-level tasks like lease abstraction at scale; and generates ranked watchlists so human attention goes where risk is highest, not where review happened to land this quarter.

The HITL design question at portfolio level is structurally identical to the deal level, just at a different scale. AI surfaces and ranks. Humans decide: which exposures get proactive outreach, which get a waiver request, which go to watch list, which trigger a workout conversation. AI-powered early warning systems have reduced non-performing loan formation by 12 to 25% across documented implementations. The AI doesn't reduce defaults by making decisions. It reduces defaults by giving humans better information earlier, while options still exist.

The structural shift this enables is from reactive to proactive portfolio management. The difference between a borrower calling with a problem and a system surfacing the problem while there's still time to act is, in practice, the difference between a workout and a write-down. Morgan Stanley has projected that AI will generate $34 billion in efficiency gains for the real estate industry by 2030. Portfolio-level automation, not just deal-level, is where much of that value accumulates.

Why Generic AI Tools Create Governance Problems Specific to CRE

Here is the specific failure mode that experienced CRE practitioners have encountered repeatedly with horizontal AI tools, and it's worth naming directly: the outputs look authoritative, but they carry no provenance. No source citation, no document page, no formula trail. In CRE, an output without a traceable source is not an output a professional can act on or defend.

This isn't an abstract concern. Parse a 200-page loan agreement through a generic large language model in a single pass, and you will occasionally get confabulated covenant terms, missed subordination provisions, and confidently stated figures that don't appear in the source document. Experienced practitioners have stopped being surprised by this. The problem is that the errors look exactly like the correct outputs. A junior analyst, or even a senior one under time pressure, will not catch the difference. The covenant that wasn't actually there, treated as binding, or the covenant that was there, missed entirely.

The examiner-readiness dimension compounds this. Regulators want the source document, the page, the formula, the calculated value, and the identity of the human who approved the disposition. A generic AI tool cannot produce that chain. A purpose-built system with citation architecture, built on retrieval-augmented generation principles and designed specifically for CRE document workflows, can.

One might argue that firms can address this by adding their own documentation layer on top of a generic tool. In practice, what firms actually do is add manual review because they don't trust the output, which produces the worst of both worlds: AI's cost without AI's efficiency gains.

The institutional knowledge problem is equally significant. CRE institutions have proprietary underwriting templates, credit policies, loan structuring standards, and deal-level judgment heuristics accumulated over years of closed transactions and managed defaults. A horizontal AI tool has no access to that institutional knowledge and cannot be calibrated to it. A purpose-built system can be trained on that body of work, making it meaningfully more accurate and relevant to the institution's specific credit culture and risk appetite.

HITL governance requires the AI to operate within defined boundaries and hand off at defined thresholds. Generic tools are not designed to have configurable approval paths or risk-proportionate escalation logic. They're designed to answer questions. The difference between a question-answering tool and a governed workflow component is the difference between a useful experiment and a production deployment.

What Good HITL Architecture Looks Like Across a CRE Institution

Three governance patterns separate deployments that achieve operational results from those that remain permanently in pilot.

The first is full traceability. Every input, output, and decision is logged. Source documents are cited. Formulas are preserved. Human approvals are timestamped. Together these elements constitute the end-to-end audit trail that examiners require. This isn't optional in a regulated industry; it's the foundation on which the entire system's credibility rests. An AI deployment that can't produce a complete audit trail on demand is not a deployment that's ready for examiner scrutiny, counterparty due diligence, or board-level accountability.

The second is parallel QA before automation. In mature deployments, AI and human reviewers work side by side on identical outputs until accuracy is validated. Then, and only then, are approval paths calibrated to that validated accuracy level. This sequence prevents the Opendoor failure mode: the human check isn't removed until the system has demonstrated that it doesn't need a check on that particular class of output. The discipline of running in parallel is what makes the eventual automation defensible.

The third is risk-proportionate approval paths. Routine extractions clear automatically. Flagged anomalies route to analysts. Credit decisions route to senior officers. The approval architecture is calibrated to the consequence of an error at each stage, not applied uniformly across all outputs. This is what compresses cycle times without cutting corners: the governance doesn't slow the low-stakes, high-volume work; it reserves deliberate review for the high-stakes, material decisions.

The decision layer must be explicit, not implicit. Which decisions are automated, which are AI-supported and human-approved, and which are human-only: this must be documented and revisited as AI accuracy improves and institutional confidence grows. Prophia's human-in-the-loop approach in lease abstraction achieves 99% accuracy with fast turnaround; the human review layer isn't slowing the process. It's what makes the accuracy claim defensible to counterparties and auditors.

Security and data governance are prerequisites, not afterthoughts. CRE loan files, borrower financials, and portfolio performance data are among the most sensitive financial information in existence. HITL architecture built on consumer-grade AI tools adapted for enterprise use is not architecture. It's improvisation.

The analyst's role in a mature HITL system looks different from what it does today, but it's not diminished. The manual data entry and document scrubbing have transferred to the AI layer. What remains for the analyst is strategic oversight, exception handling, sponsor assessment, and final decision authority: the work that requires professional judgment and carries professional accountability.

The trajectory is straightforward: as AI accuracy is validated in each workflow, approval paths can be recalibrated toward greater autonomy where confidence is established, while maintaining human ownership wherever accountability cannot be delegated. The question each institution faces is how systematically it pursues that recalibration, and whether it has the governance architecture to do so safely.

The 92% of CRE teams that have started piloting AI have answered the easy question. The 5% that have achieved their program objectives have answered the hard one: not whether to deploy AI, but how to design the decision layer around it.

Sources

  1. mckinsey.com
  2. kolena.com
  3. v7labs.com
  4. housingwire.com
  5. financialedinc.com
  6. thefractionalanalyst.com

More in AI in Commercial Real Estate