Deep dive — chart 2
Chart 2 — P3 baseline vs recent distribution
The how. Where Chart 1 shows when the P3 line rose, this chart shows how the P3 exclusion-score distribution shape moved. Mass moved right, into the ineligible_exclude lane.
P3 exclusion-score distribution — baseline (blue) vs recent (orange). Mass moved right, past the exclude threshold at score 67.
What you are looking at
- X-axis: eligibility-exclusion score (0 to 100)
- Y-axis: density of screenings at that score for P3
- Blue distribution: baseline window (first 30 days), what P3 scoring normally looks like
- Orange distribution: recent window (last 30 days), what P3 scoring looked like at Snapshot time
- Vertical line at score 67: threshold for routing to the ineligible_exclude lane
What the shape shift means
The orange distribution's mass has moved right compared to the blue distribution. Three things are happening at the same time:
- The peak (mode) rises. The most common P3 exclusion score used to sit near the further-screening range; now it sits at or above the exclude threshold.
- The right tail grows. More P3 screenings score above the exclude threshold.
- The left tail shrinks. Fewer P3 screenings land in the eligible_enroll range.
Clinical translation: the AI now systematically scores elderly-multimorbid + underrepresented candidates as less eligible than baseline — even as recruitment successfully pulled them in. The score does not distinguish between "high real-world burden but trial-appropriate" and "actual exclusion criterion triggered." The distribution shape movement is the mathematical fingerprint of that failure mode.
What KL-divergence quantifies
Kullback-Leibler divergence measures how different one distribution is from another. Values:
- KL < 0.05: the same distribution (random noise)
- KL 0.05-0.2: mild shift, monitor
- KL 0.2-0.5: significant shape shift, flag as event
- KL 0.5-1.0: substantial shape shift, high-severity event
- KL > 1.0: distributions are approaching non-overlap — the population the AI is scoring is behaviorally different from the population it was trained on
The sample's P3 KL-divergence hit 0.784 on Day 44 and 1.126 on Day 51. Both are well above the flag threshold, and the elevation persisted.
The clinical + regulatory read
Clinical read: the AI has drifted out of calibration for the elderly-multimorbid + underrepresented candidate profile. Manual investigator re-review should apply until retraining + validation.
Regulatory read: the shape movement is precisely the kind of "measurable drift on a Diversity-Plan-target subpopulation" the FDA Diversity Plan Guidance, EMA Reflection Paper on AI, and ICH E6(R3) proportionality language expect sponsors to detect and act on.
The next chart
The baseline-vs-recent chart shows how the P3 distribution moved. The lane-rate chart shows the operational consequence: which lane are P3 candidates getting routed to now vs baseline.
Chart 3 — lane rate per patient group →
$499 Snapshot. 3 business days.
Same shape-analysis for your sponsor's actual AI surface + full analysis report.
Buy $499