Deep dive — chart 2
Chart 2 — F3 baseline vs recent distribution
The how. Where Chart 1 shows when the F3 line dropped, this chart shows how the F3 readiness-risk-score distribution shape shifted. Mass moved left, out of the red_ground_urgent lane — SAFETY UNDER-flagging.
F3 readiness-risk-score distribution — baseline (blue) vs recent (orange). Mass shifted LEFT, away from the red_ground_urgent lane threshold at score 67.
What you are looking at
- X-axis: readiness-risk score (0 to 100, higher = more risky = more likely to ground)
- Y-axis: density of assessments at that score for F3
- Blue distribution: baseline window (first 30 days), what F3 scoring normally looks like
- Orange distribution: recent window (last 30 days), what F3 scoring looked like at Snapshot time
- Vertical line at score 67: threshold for routing to red_ground_urgent lane
What the shape shift means
The orange distribution's mass has moved left compared to the blue distribution. Three things are happening simultaneously:
- The peak (mode) drops. The most common F3 score used to be near the red-lane threshold; now it is well below it.
- The right tail shrinks. Fewer F3 assessments score above the red_ground_urgent threshold.
- The left tail grows. More F3 assessments score in the green_ready and yellow_watch range.
Sustainment translation: the AI now systematically scores mid-life combat assets in austere-environment operating (higher-than-baseline utilization + higher cycles-since-overhaul + more punishing theater) as LESS risky than baseline. The score does not distinguish between "veteran-airframe-in-garrison" and "veteran-airframe-in-combat-tempo." The distribution shape drift is the mathematical fingerprint of that sustainment failure mode.
What KL-divergence quantifies
Kullback-Leibler divergence measures how different one distribution is from another. Values:
- KL < 0.05: the same distribution (random noise)
- KL 0.05-0.2: mild shift, monitor
- KL 0.2-1.0: significant shape shift, flag as event
- KL > 1.0: substantial shape shift, high-severity event
- KL > 2.0: distributions are almost non-overlapping — the population the AI is scoring is behaviorally different from the population it was trained on
The sample's F3 KL-divergence hit 1.275 on Day 44, 1.628 on Day 51, and 2.101 on Day 58. That is far above threshold, and it stayed there through Day 86.
The sustainment + regulatory read
Sustainment read: the AI has drifted out of calibration for F3 assets in current operating conditions. Manual crew-chief / flight-surgeon re-review should apply until retraining + validation.
Regulatory read: the shape shift is precisely the kind of "measurable drift in performance on a specific subpopulation" that DoD AI Ethical Principle "Reliable" + NIST SP 800-53 rev5 RMF continuous-monitoring expect the program office to detect and act on.
The next chart
The baseline-vs-recent chart shows how the F3 distribution shifted. The lane-rate chart shows the operational consequence: which lane are F3 assets getting routed to now vs baseline.
Chart 3 — lane rate per asset group →
$499 Snapshot. 3 business days.
Same shape analysis for your program office's actual AI surface + full analysis report.
Buy $499