Deep dive — scenario walkthrough

The sample scenario — what the ED-triage AI was doing and why it failed silently

Mid-market community/regional hospital, ~30,000-90,000 ED encounters annually, running an AI severity classifier at the front door. Standard-of-care AI deployment. Standard operational instrumentation. Standard silent failure.

The AI system

Model: ED-triage severity classifier (mock version ed-triage-classifier-v3.1.2).

Function: at ED registration, each encounter's chief complaint + admission source + comorbidity profile + LOS estimate is scored 0-100 for severity. Score routes the encounter to one of three lanes:

  • Fast-Track (score 0-33) — low-acuity, mid-level provider
  • Standard Workup (score 34-66) — moderate-acuity, attending workup
  • Urgent Intervention (score 67-100) — high-acuity, ICU consult / trauma team

Training window: encounters through 2026-03-01. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.

The patient groups

The audit segmentation used five common ED patient groups:

Patient groupDefinition
H1Young adult, low comorbidity
H2Middle-aged, moderate comorbidity
H3Age 65+ with 3+ comorbidities (elderly multimorbid)
H4Pediatric
H5Obstetric

What shifted during the 90 days

Two input distributions shifted silently and simultaneously:

  1. Patient-mix shift. Community-outreach programs + demographic aging amplified H3 representation. H3 share of ED volume grew ~25% post-Day 45.
  2. Presentation-shape shift. Post-shift H3 encounters skewed toward atypical elderly presentations: altered mental status, non-specific abdominal pain, fever/infection without localizing signs. Exactly the pattern where a classifier trained on textbook younger-patient presentations UNDER-scores severity.
The AI was not retrained. Its baked-in miscalibration for atypical elderly presentations compounded with the shifted-input mix. H3 urgent-intervention lane rate collapsed from 43% baseline to 9% recent. A 34 percentage-point differential, in the wrong direction, silent to every aggregate metric the hospital was watching.

What the hospital's dashboards showed (green)

Standard operational instrumentation cannot see group-differential in aggregate. That is not a criticism of the operational team; it is a structural property of aggregation.

What the independent-verifier sensor saw

See how the drift-detection chart is read →

The one sentence a Chief Medical Officer would care about

The AI ED-triage engine's urgent-intervention lane rate for H3 (age 65+, 3+ comorbidities) dropped from 43% to 9% over the audit period — a 34-percentage-point group-differential in the wrong direction that occurred silently while the ED-throughput dashboard showed all-green throughout.

$499 Snapshot on your hospital's actual AI system.

Same scenario structure. Your data, your AI, your facility footprint. Same 3 business day turnaround.

Buy $499