Deep dive — scenario walkthrough
The sample scenario — what the mortgage-underwriting AI was doing and why it failed silently
Mid-market mortgage lender, ~$2B-$10B annual originations (mid-market bank OR non-bank independent mortgage banker), running an AI underwriting classifier at the front of the pipeline. Standard-of-industry AI deployment. Standard operational instrumentation. Standard silent failure.
The AI system
Model: mortgage-underwriting classifier (mock version mortgage-underwriter-v4.7.2).
Function: at application intake, each file's stated purpose + credit-bureau snapshot + property valuation + LTV + DTI + FICO tier is scored 0-100. Score routes the file to one of three lanes:
- Auto-Approve Conforming (score 0-33) — strong file, GSE-conforming auto-approve
- Manual Review + Upcharge (score 34-66) — tier-adjusted rate, manual verification
- Decline / Refer non-QM (score 67-100) — decline OR refer to non-QM lender at higher rate
Training window: applications through 2026-01-15. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.
The applicant groups
The audit segmentation used five common mortgage-applicant groups:
| Applicant group | Definition |
| H1 | Prime credit, majority-white zip, low LTV |
| H2 | Prime credit, mixed zip, moderate LTV |
| H3 | Age 30-50, near-prime credit, majority-minority zip |
| H4 | First-time homebuyer, prime credit |
| H5 | Self-employed, non-W2 income |
What shifted during the 90 days
Two input distributions shifted silently and simultaneously:
- Applicant-mix shift. Marketing outreach amplified Group H3 representation. H3 share of application volume grew ~30% post-Day 45.
- File-attribute shift. Post-shift H3 files skewed toward FHA (higher-LTV product) + LTV +14pts + DTI +8pts + FICO -32pts. Underwriter attributes moved into the near-prime-slipping-to-sub-prime edge.
The AI was not retrained. Its baked-in H3-adverse offset (+5 underwriting-risk-score points, from historical training) compounded with the shifted-input mix. H3 decline / non-QM lane rate rose from 14% baseline to 61% recent. A 47-percentage-point differential, in the wrong direction, silent to every aggregate metric the lender was watching.
What the lender's dashboards showed (green)
- Origination throughput — stable
- Mean underwriting-risk-score — stable
- Aggregate approval rate — stable (H1/H2/H4/H5 held, H3 collapsed, aggregate rounded to noise)
- Model confidence — stable
- Loan-loss ratio — stable (30-60-90 day lag masks any credit-quality signal)
- Quarterly board risk report — all-green
Standard operational instrumentation cannot see group-differential in aggregate. That is not a criticism of the operational team; it is a structural property of aggregation.
What the independent-verifier sensor saw
- H3 score distribution shape diverged from baseline (KL > 0.8) starting Day 44 and never recovered
- H3 decline / non-QM routing rate rose sharply across every 7-day window from Day 44 forward
- H5 self-employed distribution also diverged (Days 30, 37, 79, 86) — second signal on a separate group
- 11 total drift events, 10 high-severity, 1 medium-severity
See how the drift-detection chart is read →
The one sentence a Chief Compliance Officer would care about
The AI mortgage-underwriting engine's decline / non-QM lane rate for Group H3 (age 30-50, near-prime credit, majority-minority zip) rose from 14% to 61% over the audit period — a 47-percentage-point group-differential in the wrong direction that ran silently while the lender's origination-dashboard, aggregate approval rate, and loan-loss ratio all showed all-green throughout.
$499 Snapshot on your lender's actual AI system.
Same scenario structure. Your data, your AI, your state / charter footprint. Same 3 business day turnaround.
Buy $499