{
  "abstract": "Background Colorectal cancer prevention relies on accurate polyp detection during colonoscopy, yet clinical translation of deep learning segmentation demands real-time inference, a constraint standard architectures struggle to meet on resource-limited endoscopy hardware. Whether lightweight architectures can close the accuracy gap while dramatically reducing computational cost remains an open question with direct clinical implications. We benchmarked three encoder-decoder architectures to establish practical feasibility for live colonoscopy deployment.Methods Three architectures were evaluated on Kvasir-SEG (1,000 expert-annotated images; 880/60/60 train/val/test split): U-Net with ResNet34 (24.4M parameters), DeepLabV3+ with ResNet50 (26.7M parameters), and a lightweight U-Net with MobileNetV2 (6.6M parameters - 4× reduction). All used ­ImageNet-pretrained encoders were trained for 30 epochs with a combined Dice-BCE loss, Adam optimisation (lr=1e-4), and learning rate scheduling. Inference latency was benchmarked on an NVIDIA A100 GPU at 256×256 resolution. Per-image Dice scores were compared using paired t-tests.Results All three models converged stably with no overfitting ( IDDF2026-ABS-0264 Figure 3. Training and validation loss curves across 30 epochs for all three architectures), with progressive Dice (IDDF2026-ABS-0264 Figure 4. Training and validation dice score trajectories across 30 epochs) and IoU (IDDF2026-ABS-0264 Figure 5. Training and validation IoU trajectories across 30 epochs) improvements across 30 epochs. The accuracy-speed trade-off across architectures is summarised in (IDDF2026-ABS-0264 Figure 1. Accuracy vs inference speed trade-off plot for all three architectures). DeepLabV3+ achieved the highest Dice (0.880±0.150; IoU=0.811) at 123.0 FPS. U-Net ResNet34 reached Dice 0.867±0.156 at 154.6 FPS. The lightweight MobileNetV2 U-Net achieved Dice 0.851±0.171 at 130.9 FPS - a 4.4× real-time margin above the 30 FPS clinical threshold, using only 25% of the parameters of standard models (IDDF2026-ABS-0264 Figure 1. Accuracy vs inference speed trade-off plot for all three architectures). Per-image Dice distributions confirmed comparable score distributions across all three architectures, with overlapping interquartile ranges (IDDF2026-ABS-0264 Figure 2. Per-image dice score distribution on the test set across all three architectures). Critically, all three architectures exceeded the real-time threshold by 4-5×, providing headroom for higher resolutions, additional post-processing, and lower-specification clinical hardware.Conclusions A MobileNetV2 U-Net with 4× fewer parameters achieves near-equivalent segmentation accuracy (Dice 0.851 vs 0.880) with a 4.4× real-time margin, making ­lightweight ­architectures robust candidates for deployment in GPU-­accelerated endoscopy systems. This computational headroom supports operation at higher resolutions and on mid-range clinical hardware without breaching the real-time boundary. These findings support a shift toward parameter-efficient architectures for clinical polyp detection. Multi-centre validation and edge-device benchmarking are warranted as next steps.Abstract IDDF2026-ABS-0264 Figure 1Abstract IDDF2026-ABS-0264 Figure 2Abstract IDDF2026-ABS-0264 Figure 3Abstract IDDF2026-ABS-0264 Figure 4Abstract IDDF2026-ABS-0264 Figure 5",
  "authors": [
    {
      "affiliations": [
        "New York Medical College, United States"
      ],
      "name": "Jeril Lasington"
    },
    {
      "affiliations": [
        "Rutgers University, United States"
      ],
      "name": "Lawin Steve Mathew Lasington"
    },
    {
      "affiliations": [
        "Boston University, United States"
      ],
      "name": "Swamynathan Umamaheshwaran"
    }
  ],
  "title": "IDDF2026-ABS-0264 Small model, big impact: can lightweight deep learning match standard architectures for real-time polyp segmentation?",
  "uid": "df203ad4-9884-5e42-a430-361b7985581f"
}
