Deep dive — scenario walkthrough
The sample scenario — what the predictive-maintenance AI was doing and why it under-flagged silently
Defense-industrial-base program office / prime contractor, fleet size 500-5,000 assets, multi-theater operating footprint, running an AI predictive-maintenance classifier on every mission-cycle. Standard sustainment-AI deployment. Standard fleet-readiness instrumentation. Standard silent failure — in the SAFETY-critical direction.
The AI system
Model: predictive-maintenance classifier (mock version predictive-maintenance-classifier-v7.2.3).
Function: at each mission-cycle, an asset's utilization + mission-type + operating-theater + cycles-since-overhaul + repair-cost estimate is scored 0-100 for readiness-risk. Score routes the asset to one of three lanes:
- Green -- Mission Capable (score 0-33) — no scheduled action
- Yellow -- Watch / Scheduled (score 34-66) — scheduled preventative window
- Red -- Ground Urgent (score 67-100) — ground for urgent repair / mission-critical component replacement
Training window: asset data through 2025-12-31. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.
The asset groups
The audit segmentation used five common fleet asset groups:
| Asset group | Definition |
| F1 | New airframe / low cycles-since-overhaul, garrison-based |
| F2 | Mid-life airframe, standard utilization, garrison-based |
| F3 | Mid-life combat assets, mixed-utilization, austere-environment operating (CENTCOM / AFRICOM / INDOPACOM forward-deployed) |
| F4 | Late-life airframe, high cycles-since-overhaul |
| F5 | Training / test-and-evaluation assets |
What shifted during the 90 days
Two input distributions moved silently and simultaneously:
- Utilization-mix shift. Operational tempo increased F3 representation. F3 share of the fleet-readiness stream grew ~30% post-Day 45.
- Mission + theater shift. Post-shift F3 assets skewed toward combat_operations + surveillance_ISR mission types + austere theaters (CENTCOM / AFRICOM / INDOPACOM forward-deployed) + higher cycles-since-overhaul (accelerated wear from op-tempo). Exactly the pattern where a classifier trained on F3-as-reliable-veteran-airframe UNDER-scores readiness-risk.
The AI was not retrained. Its baked-in F3-as-reliable prior compounded with the shifted-input mix. F3 red_ground_urgent lane rate collapsed from 34% baseline to 8% recent. A 26 percentage-point differential, in the WRONG direction (SAFETY UNDER-flagging), silent to every aggregate metric the program office was watching.
What the program office's dashboards showed (green)
- Fleet-wide assessment throughput — stable
- Mean readiness-risk score — stable
- Aggregate red-lane count — stable (F1/F2/F4/F5 held, F3 collapsed, aggregate rounded to noise)
- Model confidence — stable
- Aggregate maintenance backlog — stable
Standard fleet-readiness instrumentation cannot see group-differential in aggregate. That is not a criticism of the sustainment team; it is a structural property of aggregation.
What the independent-verifier sensor saw
- F3 score distribution shape diverged from baseline (KL-divergence > 1.0) starting Day 44
- F3 red_ground_urgent routing rate dropped ~26 percentage points across every 7-day window from Day 44 forward
- Similar (weaker) F4 divergence started Day 51 — late-life airframes also drifting
- 13 total drift events, 12 high-severity, 1 medium-severity
See how the drift-detection chart is read →
The one sentence a Program Manager and Chief Engineer would care about
The AI predictive-maintenance classifier's red_ground_urgent lane rate for Asset Group F3 (mid-life combat assets, mixed-utilization, austere-environment operating) collapsed from 34% to 8% over the audit period — a 26 percentage-point group-differential in the WRONG direction (SAFETY UNDER-flagging) that occurred silently while the aggregate fleet-readiness dashboard showed all-green throughout.
Why UNDER-flagging is the harder case
Over-flagging is expensive: unnecessary groundings + inflated sustainment cost + operator irritation. Everyone notices. Under-flagging is silent by construction — a plane that should have been grounded and was not is only visible after the fact, and only if the fact happens. That is the DoD 3000.09 / DoD AI Ethical Principle "Reliable" exposure profile, and it is why the sensor is built to catch drift in EITHER direction.
$499 Snapshot on your program office's actual AI system.
Same scenario structure. Your data, your AI, your ATO scope. Same 3 business day turnaround. No classified / CUI / ITAR data required — anonymized readiness-assessment records are sufficient.
Buy $499