Deep dive — scenario walkthrough
The sample scenario — what the trial-eligibility AI was doing and why it failed silently
Mid-cap biopharma sponsor ($500M-$5B market cap), running an oncology Phase III trial with an AI-assisted eligibility classifier at the screening desk. Standard clinical-AI deployment. Standard CTMS + EDC + eTMF instrumentation. Standard silent failure.
The AI system
Model: trial-eligibility classifier (mock version trial-eligibility-classifier-v2.8.1).
Function: at screening, each candidate's primary diagnosis + ECOG performance status + prior-therapy lines + comorbidity count + days-to-randomization-target is scored 0-100 for exclusion likelihood. Score routes the candidate to one of three lanes:
- Eligible / Enroll (score 0-33) — clear inclusion criteria met, auto-eligible
- Further Screening / Investigator Review (score 34-66) — borderline, manual review
- Ineligible / Exclude (score 67-100) — exclusion criteria triggered
Training window: historical trial data through 2026-01-15. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.
The patient groups
The audit segmentation used five common oncology-trial patient groups:
| Patient group | Definition |
| P1 | Young adult, low comorbidity, first-line eligible |
| P2 | Middle-aged, moderate comorbidity, 1-2 prior therapy lines |
| P3 | Age 65+, 3+ comorbidities, minority representation (elderly multimorbid + underrepresented) |
| P4 | Pediatric / adolescent + young adult (AYA) |
| P5 | Prior-line-heavy (3+ prior therapy lines) refractory |
What shifted during the 90 days
Two input distributions shifted silently and at the same time:
- Recruitment-mix shift. Community-partnership + referral-network outreach amplified P3 share by ~35% post-Day 45 as the sponsor executed the trial's Diversity Plan commitment.
- Real-world-burden shift. Post-shift P3 candidates presented with higher comorbidity counts, worse ECOG performance status, and more prior-therapy lines than the P3 training distribution — exactly the real-world burden the FDA Diversity Plan Guidance instructs sponsors to include, and exactly the shape the historically-underrepresented population carries.
The AI was not retrained. Its baked-in P3-exclusion offset (+3 exclusion-score points, from historical P3-underrepresented training) stacked with the shifted-input mix. P3 candidates were routed to the ineligible_exclude lane at 75% frequency vs the 22% baseline. A 53 percentage-point differential, silent to every aggregate metric the sponsor was watching.
What the sponsor's dashboards showed (green)
- CTMS enrollment velocity — stable
- Mean eligibility-exclusion score — stable
- Aggregate screen-fail rate — stable (P1/P2/P4 held, P3 collapsed toward exclude, aggregate rounded to noise)
- Model confidence — stable
- Site-level enrollment funnel — stable
Standard clinical-operations instrumentation cannot see patient-group-level differential in the aggregate. That is not a criticism of the ClinOps team; it is a structural property of aggregation.
What the independent-verifier sensor saw
- P3 score distribution shape diverged from baseline (KL-divergence 0.784) starting Day 44
- P3 exclude-lane routing rate climbed steadily from Day 44 forward, peaking at 75% by the recent window
- P3 KL-divergence hit 1.126 on Day 51 — 5x the flag threshold
- 15 total drift events, 10 high-severity, 5 medium-severity
See how the drift-detection chart is read →
The one sentence a VP of Regulatory Affairs would care about
The AI trial-eligibility classifier's ineligible_exclude lane rate for P3 (age 65+, 3+ comorbidities, minority representation) rose from 22% to 75% over the audit period — a 53 percentage-point differential that occurred silently while the sponsor's CTMS enrollment-velocity dashboard, mean-eligibility-score view, and aggregate screen-fail-rate report all stayed in normal range throughout.
$499 Snapshot on your sponsor's actual AI system.
Same scenario structure. Your data, your AI, your trial phase + indication. Same 3 business day turnaround.
Buy $499