Deep dive — scenario walkthrough

The sample scenario — what the trial-eligibility AI was doing and why it failed silently

Mid-cap biopharma sponsor ($500M-$5B market cap), running an oncology Phase III trial with an AI-assisted eligibility classifier at the screening desk. Standard clinical-AI deployment. Standard CTMS + EDC + eTMF instrumentation. Standard silent failure.

The AI system

Model: trial-eligibility classifier (mock version trial-eligibility-classifier-v2.8.1).

Function: at screening, each candidate's primary diagnosis + ECOG performance status + prior-therapy lines + comorbidity count + days-to-randomization-target is scored 0-100 for exclusion likelihood. Score routes the candidate to one of three lanes:

  • Eligible / Enroll (score 0-33) — clear inclusion criteria met, auto-eligible
  • Further Screening / Investigator Review (score 34-66) — borderline, manual review
  • Ineligible / Exclude (score 67-100) — exclusion criteria triggered

Training window: historical trial data through 2026-01-15. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.

The patient groups

The audit segmentation used five common oncology-trial patient groups:

Patient groupDefinition
P1Young adult, low comorbidity, first-line eligible
P2Middle-aged, moderate comorbidity, 1-2 prior therapy lines
P3Age 65+, 3+ comorbidities, minority representation (elderly multimorbid + underrepresented)
P4Pediatric / adolescent + young adult (AYA)
P5Prior-line-heavy (3+ prior therapy lines) refractory

What shifted during the 90 days

Two input distributions shifted silently and at the same time:

  1. Recruitment-mix shift. Community-partnership + referral-network outreach amplified P3 share by ~35% post-Day 45 as the sponsor executed the trial's Diversity Plan commitment.
  2. Real-world-burden shift. Post-shift P3 candidates presented with higher comorbidity counts, worse ECOG performance status, and more prior-therapy lines than the P3 training distribution — exactly the real-world burden the FDA Diversity Plan Guidance instructs sponsors to include, and exactly the shape the historically-underrepresented population carries.
The AI was not retrained. Its baked-in P3-exclusion offset (+3 exclusion-score points, from historical P3-underrepresented training) stacked with the shifted-input mix. P3 candidates were routed to the ineligible_exclude lane at 75% frequency vs the 22% baseline. A 53 percentage-point differential, silent to every aggregate metric the sponsor was watching.

What the sponsor's dashboards showed (green)

Standard clinical-operations instrumentation cannot see patient-group-level differential in the aggregate. That is not a criticism of the ClinOps team; it is a structural property of aggregation.

What the independent-verifier sensor saw

See how the drift-detection chart is read →

The one sentence a VP of Regulatory Affairs would care about

The AI trial-eligibility classifier's ineligible_exclude lane rate for P3 (age 65+, 3+ comorbidities, minority representation) rose from 22% to 75% over the audit period — a 53 percentage-point differential that occurred silently while the sponsor's CTMS enrollment-velocity dashboard, mean-eligibility-score view, and aggregate screen-fail-rate report all stayed in normal range throughout.

$499 Snapshot on your sponsor's actual AI system.

Same scenario structure. Your data, your AI, your trial phase + indication. Same 3 business day turnaround.

Buy $499