{
  "abstract": "Background Endoscopic Mayo scoring in UC carries inter-observer variability of ~30%, limiting treat-to-target monitoring reliability. Deep learning models offer consistent automated scoring, but safe clinical deployment requires not only accuracy but well-calibrated confidence estimates where stated certainty reflects true likelihood of correctness. We report accuracy, calibration, and a confidence-based human review flagging system from a validated UC endoscopy AI.Methods An EfficientNet-B3 classifier was trained on 9,590 colonoscopy images (LIMUC dataset, 564 patients) to predict Mayo endoscopic grade (0-3). Performance on a held-out test set (n=1,686) was assessed using quadratic weighted kappa (QWK), MAE, and macro F1. Calibration was quantified via Expected Calibration Error (ECE, 10-bin) and per-class Brier scores. A confidence-based flagging system identified high-uncertainty predictions (entropy >Q75) for mandatory human review.Results Training converged at epoch 4 (best validation QWK=0.840), with stable generalisation thereafter despite continued training improvement ( IDDF2026-ABS-0255 Figure 1. Model training curve). The model achieved QWK=0.840, MAE=0.242, and macro F1=0.723, exceeding published inter-human QWK (~0.60). The confusion matrix confirmed the strongest performance at diagnostic extremes (Mayo 0: 730/925 correct; Mayo 3: 75/120 correct), with expected confusion at intermediate grades (IDDF2026-ABS-0255 Figure 3. Confusion matrix for four class mayo grade prediction).Calibration was excellent at ECE=0.024 well below estimated human expert consistency (~0.08) and typical uncalibrated CNNs (~0.15) without requiring post-hoc temperature scaling. Per-class Brier scores were lowest for severe disease (Mayo 3=0.030) and highest at boundary grades (Mayo 1=0.130). Entropy-based flagging (IDDF2026-ABS-0255 Figure 4. Uncertainty distribution by Mayo grade) identified 25% of cases (n=422) as high-uncertainty, with flagged case accuracy of 51.9% versus 80.5% in confident predictions (IDDF2026-ABS-0255 Figure 2. AI safety net effect) confirming the safety net effect. Mayo 1 and Mayo 2 were disproportionately flagged (35.3% and 30.5% respectively) (IDDF2026-ABS-0255 Figure 5. Disagreement flagging rate by Mayo grade), precisely the grades with the highest mean absolute prediction error (Mayo 2=0.379, Mayo 3=0.400) (IDDF2026-ABS-0255 Figure 6. Mean absolute grade error by disease severity).Conclusions This UC endoscopic AI achieves expert-level scoring agreement (QWK=0.840) with well-calibrated confidence (ECE=0.024) and a principled human review flagging system. The combination of high accuracy, trustworthy uncertainty quantification, and targeted boundary-grade flagging addresses the three core prerequisites for safe clinical deployment of AI endoscopic scoring in UC.Abstract IDDF2026-ABS-0255 Figure 1Abstract IDDF2026-ABS-0255 Figure 2Abstract IDDF2026-ABS-0255 Figure 3Abstract IDDF2026-ABS-0255 Figure 4Abstract IDDF2026-ABS-0255 Figure 5Abstract IDDF2026-ABS-0255 Figure 6",
  "authors": [
    {
      "affiliations": [
        "Rutgers University, United States"
      ],
      "name": "Lawin Steve Mathew Lasington"
    },
    {
      "affiliations": [
        "New York Medical College, United States"
      ],
      "name": "Jeril Lasington"
    },
    {
      "affiliations": [
        "Boston University, United States"
      ],
      "name": "Swamynathan Umamaheshwaran"
    }
  ],
  "title": "IDDF2026-ABS-0255 Ai-assisted mayo endoscopic scoring in ulcerative colitis: accuracy, calibration, and confidence-based flagging for safe clinical integration",
  "uid": "d04c4f76-df1a-5ab5-8946-14878695b976"
}
