Deep dive — scenario walkthrough
The sample scenario — what the AI credit-underwriting engine was doing and why it failed silently
Mid-market consumer bank, $5B-$50B assets, 2,000-20,000 consumer credit applications monthly, running an AI classifier at the front door of every personal-loan, credit-card, auto-loan, and HELOC application. Standard-of-market AI deployment. Standard model-risk instrumentation. Standard silent failure.
The AI system
Model: consumer-credit-underwriting classifier (mock version credit-underwriter-v6.1.4).
Function: at intake, each application's stated purpose + credit-bureau snapshot at decision-time + DTI + requested amount is scored 0-100 for underwriting risk. Score routes the application to one of three lanes:
- Auto-Approve Prime (score 0-33) — strong profile, best-APR auto-approval
- Manual Review / Price-Up (score 34-66) — tier-adjusted APR, manual verification
- Decline / Secondary Market (score 67-100) — declined or referred to secondary-market / higher-APR product
Training window: applicant data through 2026-01-31. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.
The borrower groups
The audit segmentation used five common consumer-credit borrower groups:
| Borrower group | Definition |
| B1 | Prime, established credit, higher-income zip |
| B2 | Near-prime, moderate credit, mixed zip |
| B3 | Age 25-45, near-prime credit, mixed lower-income zip |
| B4 | Sub-prime, thin file, mixed zip |
| B5 | Sub-prime, established derogatory, lower-income zip |
What shifted during the 90 days
Two input distributions shifted silently and simultaneously:
- Applicant-mix shift. Marketing outreach into mixed lower-income zips amplified B3 representation. B3 share of application volume grew ~30% post-Day 45.
- Applicant-profile shift. Post-shift B3 applicants skewed toward stated_purpose=debt_consolidation, higher DTI, larger requested amounts, and modestly weaker FICO-at-decision. Marketing surfaced debt-management prospects rather than the pre-shift prime-mix B3 pool. Exactly the profile where a classifier trained on the pre-shift mix systematically over-scores underwriting risk.
The AI was not retrained. Its baked-in offset for the pre-shift B3 profile compounded with the shifted-input mix. B3 decline / secondary-market lane rate rose from 18% baseline to 62% recent. A 44 percentage-point differential, in the wrong direction, silent to every aggregate model-risk metric the bank was watching.
What the bank's model-risk dashboards showed (green)
- Application throughput — stable
- Mean underwriting-risk score — stable
- Aggregate approval rate — stable (B1 / B4 / B5 held, B2 modest shift, B3 decline surge, aggregate rounded to noise)
- Portfolio loss ratio — stable (no time yet for the misroutes to season into charge-offs)
- Model confidence — stable
Standard model-risk instrumentation cannot see group-differential in aggregate. That is not a criticism of the model-risk team; it is a structural property of aggregation.
What the independent-verifier sensor saw
- B3 score distribution shape diverged from baseline (KL-divergence > 0.5) starting Day 44
- B3 decline / secondary-market routing rate rose steadily across every 7-day window from Day 44 forward
- Repeated per-group KL-divergence signals across B2, B3, B5 through the audit period
- 20 total drift events, 14 high-severity, 6 medium-severity
See how the drift-detection chart is read →
The one sentence a Chief Compliance Officer would care about
The AI credit-underwriting engine's decline / secondary-market lane rate for B3 (age 25-45, near-prime credit, mixed lower-income zip) rose from 18% to 62% over the audit period — a 44 percentage-point group-differential in the wrong direction that occurred silently while the aggregate approval rate and portfolio loss ratios stayed within tolerance throughout.
$499 Snapshot on your bank's actual AI system.
Same scenario structure. Your data, your AI, your charter footprint. Same 3 business day turnaround.
Buy $499