Deep dive — scenario walkthrough
The sample scenario — what the admissions AI was doing and why it failed silently
R1 or R2 research university, $1B-$5B endowment, 25,000-65,000 applications per admission cycle, running an AI decision-support classifier that scores every applicant and routes them to a lane. Standard vendor-stack deployment. Standard operational dashboards. Standard silent failure.
The AI system
Model: admissions-recommender classifier (mock version admissions-recommender-v3.4.1).
Function: at application intake, each applicant file's academic profile + intended major + financial-need estimate + high-school context + extracurricular record is scored 0-100. Score routes the application to one of three lanes:
- Auto-Admit (score 0-33) — strong applicant, auto-admit recommendation
- Committee Review (score 34-66) — borderline applicant, requires human committee review
- Deny / Waitlist (score 67-100) — deny or waitlist recommendation
Training window: applications through 2026-02-01. Deployed Day 1 of the audit period. Not retrained during the 90-day production run.
The applicant groups
The audit segmentation used five common admissions applicant groups:
| Applicant group | Definition |
| A1 | Full-pay, well-resourced private + suburban public school applicants |
| A2 | Mid-income, mixed-resource school applicants |
| A3 | First-generation, mid-low-income, under-resourced public school applicants |
| A4 | International applicants |
| A5 | Transfer + non-traditional applicants |
What shifted during the 90 days
Two input distributions shifted silently and simultaneously:
- Applicant-mix shift. Recruiter outreach into under-resourced public-school markets amplified A3 representation. A3 share of applicant volume grew ~30% post-Day 45.
- Applicant-profile shift. Post-shift A3 applicants skewed toward humanities and social-sciences intended majors, higher estimated financial need, and modestly lower standardized-test scores and GPA. Exactly the pattern where a classifier trained on a historically A3-thin distribution over-penalizes.
The AI was not retrained. Its baked-in socioeconomic-proxy penalty (+9 recommendation-score points for A3 applicants) compounded with the shifted-input mix. A3 deny-or-waitlist lane rate rose from 23% baseline to 68% recent. A 45 percentage-point differential, in the wrong direction, silent to every aggregate metric the institution was watching.
What the institution's dashboards showed (green)
- Total applications received — up (recruiter outreach worked)
- Aggregate admission rate — stable
- Aggregate yield rate — stable
- Aggregate diversity dashboard with post-hoc adjustments — stable
- Model confidence — stable
Standard enrollment-management instrumentation cannot see group-differential in aggregate. That is not a criticism of the enrollment team; it is a structural property of aggregation.
What the independent-verifier sensor saw
- A3 score distribution shape diverged from baseline (KL-divergence up to 0.727) starting Day 44
- A3 deny-or-waitlist lane rate rose sharply and stayed up across every 7-day window from Day 44 forward
- Cross-applicant-group differential signal fired repeatedly, primary source A3
- 12 total drift events, 10 high-severity, 2 medium-severity
See how the drift-detection chart is read →
The one sentence a General Counsel would care about
The AI admissions-recommender's deny-or-waitlist lane rate for A3 (first-generation, mid-low-income, under-resourced public school applicants) rose from 23% to 68% over the audit period — a 45 percentage-point group-differential in the wrong direction that occurred silently while the aggregate admission-rate and yield-rate dashboard showed all-green throughout.
$499 Snapshot on your institution's actual AI system.
Same scenario structure. Your data, your admissions AI, your accreditor footprint. Same 3 business day turnaround.
Buy $499