{
  "abstract": "Background Aggregate accuracy metrics obscure clinically dangerous AI failures at critical Mayo Endoscopic Score (MES) boundaries. The Mayo 1 - 2 boundary governing treat-or-watch decision represents the highest-stakes classification in UC management, yet its specific failure profile remains uncharacterised in published AI systems. We performed a systematic subgroup failure analysis to identify where AI endoscopic scoring poses a genuine clinical risk.Methods An EfficientNet-B3 model was trained on the LIMUC dataset (9,590 images, 564 patients) and evaluated on a held-out test set (n=1,686). Beyond aggregate metrics, boundary-level failure analysis was performed at all three clinical decision thresholds (Mayo 0 - 1, 1 - 2, 2 - 3), quantifying under-grading (under-treatment risk) and over-grading (unnecessary escalation risk) rates separately. Prediction entropy identified a high-risk zone combining boundary ambiguity with model uncertainty.Results Despite strong overall performance (QWK=0.84, accuracy=77.2%), subgroup analysis revealed progressive confidence loss across severity grades in Mayo 0: 74.5% ­confidence versus Mayo 2: 54.8% (IDDF2026-ABS-0256 ­Figure 1. Per-grade ­accuracy vs model confidence across all four Mayo grades). Boundary-level error rates confirmed clinically dangerous failure profiles at all three decision thresholds (IDDF2026-ABS-0256 Figure 2. Boundary level error rates at all three clinical decision thresholds): Mayo 0 - 1 under-grading 17.4%, over-grading 17.7%; Mayo 1 - 2 under-grading 17.1%, over-grading 12.2%; Mayo 2 - 3 showing the highest under-grading rate of 28.8%, meaning nearly 1 in 3 severe cases would be under-escalated.Uncertainty analysis revealed Mayo 1 and Mayo 2 as the highest-entropy grades, with Mayo 1 - 2 and Mayo 2 - 3 boundaries showing the greatest prediction uncertainty (IDDF2026-ABS-0256 Figure 3. Model uncertainty prediction entropy distribution by Mayo grade). A high-risk zone combining Mayo 1|2 cases with elevated prediction entropy (n=268) yielded accuracy of only 69.0% versus 74.5% in confident cases (IDDF2026-ABS-0256 Figure 4. Accuracy collapses in the high-risk zone), identifying a clinically actionable subgroup requiring mandatory human override.Conclusions Aggregate AI accuracy of 77% conceals boundary-specific failure rates of 17-29% at clinically consequential decision thresholds. Safe UC AI integration requires subgroup-aware validation, mandatory clinician override at high-uncertainty boundaries, and reporting standards that disclose grade-level performance. These findings provide an evidence base for regulatory frameworks governing AI deployment in IBD endoscopy.Abstract IDDF2026-ABS-0256 Figure 1Abstract IDDF2026-ABS-0256 Figure 2Abstract IDDF2026-ABS-0256 Figure 3Abstract IDDF2026-ABS-0256 Figure 4",
  "authors": [
    {
      "affiliations": [
        "New York Medical College, United States"
      ],
      "name": "Jeril Lasington"
    },
    {
      "affiliations": [
        "Rutgers University, United States"
      ],
      "name": "Lawin Steve Mathew Lasington"
    },
    {
      "affiliations": [
        "Boston University, United States"
      ],
      "name": "Swamynathan Umamaheshwaran"
    }
  ],
  "title": "IDDF2026-ABS-0256 Identifying ai limitations in UC endoscopic scoring to ensure safe clinical integration: a subgroup failure analysis",
  "uid": "ba50403e-eb0b-596f-b3cc-2adbbff960ba"
}
