Deep dive — scenario walkthrough
The sample scenario — what the ED-triage AI was doing and why it failed silently
Mid-market community/regional hospital, ~30,000-90,000 ED encounters annually, running an AI severity classifier at the front door. Standard-of-care AI deployment. Standard operational instrumentation. Standard silent failure.
The AI system
Model: ED-triage severity classifier (mock version ed-triage-classifier-v3.1.2).
Function: at ED registration, each encounter's chief complaint + admission source + comorbidity profile + LOS estimate is scored 0-100 for severity. Score routes the encounter to one of three lanes:
- Fast-Track (score 0-33) — low-acuity, mid-level provider
- Standard Workup (score 34-66) — moderate-acuity, attending workup
- Urgent Intervention (score 67-100) — high-acuity, ICU consult / trauma team
Training window: encounters through 2026-03-01. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.
The patient groups
The audit segmentation used five common ED patient groups:
| Patient group | Definition |
| H1 | Young adult, low comorbidity |
| H2 | Middle-aged, moderate comorbidity |
| H3 | Age 65+ with 3+ comorbidities (elderly multimorbid) |
| H4 | Pediatric |
| H5 | Obstetric |
What shifted during the 90 days
Two input distributions shifted silently and simultaneously:
- Patient-mix shift. Community-outreach programs + demographic aging amplified H3 representation. H3 share of ED volume grew ~25% post-Day 45.
- Presentation-shape shift. Post-shift H3 encounters skewed toward atypical elderly presentations: altered mental status, non-specific abdominal pain, fever/infection without localizing signs. Exactly the pattern where a classifier trained on textbook younger-patient presentations UNDER-scores severity.
The AI was not retrained. Its baked-in miscalibration for atypical elderly presentations compounded with the shifted-input mix. H3 urgent-intervention lane rate collapsed from 43% baseline to 9% recent. A 34 percentage-point differential, in the wrong direction, silent to every aggregate metric the hospital was watching.
What the hospital's dashboards showed (green)
- Door-to-doc time — stable
- Mean triage score — stable
- Aggregate urgent-intervention lane count — stable (H1/H2/H4/H5 held, H3 collapsed, aggregate rounded to noise)
- Model confidence — stable
- Left-without-being-seen rate — stable
Standard operational instrumentation cannot see group-differential in aggregate. That is not a criticism of the operational team; it is a structural property of aggregation.
What the independent-verifier sensor saw
- H3 score distribution shape diverged from baseline (KL-divergence > 2.0) starting Day 44
- H3 fast_track routing rate increased +11-14% across every 7-day window from Day 44 forward
- Cross-patient group differential signal fired Day 89 (fast_track shift +11.3% for H3 vs -1.1% for H1)
- 17 total drift events, 9 high-severity, 8 medium-severity
See how the drift-detection chart is read →
The one sentence a Chief Medical Officer would care about
The AI ED-triage engine's urgent-intervention lane rate for H3 (age 65+, 3+ comorbidities) dropped from 43% to 9% over the audit period — a 34-percentage-point group-differential in the wrong direction that occurred silently while the ED-throughput dashboard showed all-green throughout.
$499 Snapshot on your hospital's actual AI system.
Same scenario structure. Your data, your AI, your facility footprint. Same 3 business day turnaround.
Buy $499