{
  "abstract": "Introduction Large language models (LLMs) are increasingly being adopted across healthcare disciplines to deliver patient-oriented information. However, no systematic review has examined the accuracy and readability of LLM-generated content specific to gastrointestinal (GI) procedures, representing a critical gap in understanding the reliability of these emerging technologies for GI-related patient and clinician education.Methods We evaluated the accuracy and readability of LLM-generated responses to patient- and clinician-oriented queries regarding GI investigations and interventions, and identified barriers to clinical implementation. A systematic search of Ovid MEDLINE and Ovid Embase identified peer-reviewed literature from January 2023 to July 2025 examining text-based LLM outputs related to GI investigative procedures (endoscopy, colonoscopy) and surgical interventions (cholecystectomy, colectomy, gastrectomy, pancreatic resection). The NIH Quality Assessment Tool for observational cohort studies was used to assess risk of bias. Accuracy scores were reported with n representing the number of questions evaluated.Results Of the 432 screened articles, 24 underwent full-text review, with 11 meeting inclusion criteria. Accuracy varied by model and scoring system. On binary scales, Google Bard demonstrated the highest accuracy (85.7%, n=17 questions), followed by ChatGPT (71.4%, n=17). On a three-point scale, Claude performed best (2.68/3, n=28), followed by Mixtral (2.3/3, n=28) and ChatGPT (2.24/3, n=74). Using a five-point scale, ChatGPT outperformed other models (3.57/5, n=261). By procedure type, endoscopy achieved the highest scores on five-point scales (4.81/5, n=7) and second-highest on three-point scales (2.69/3, n=9), while colonoscopy scored highest on three-point scales (2.78/3, n=8) and second-highest on five-point scales (4.28/5, n=13). Among surgical procedures, open pancreatic head resection scored highest (4.40/5, n=20), followed by open gastrectomy (4.35/5, n=20). For readability, Bard achieved the highest Flesch Reading Ease score (56.3, n=66), followed by ChatGPT-4 (42.7, n=66) and ChatGPT-3 (32.86, n=6).Conclusions While LLMs can generate readable content on GI procedures, none achieve the ≥95% accuracy threshold typically expected for clinical use, with performance remaining inconsistent across procedures and models. Standardised benchmarking protocols, balanced prompt libraries, and ongoing model evaluations are necessary before LLMs can be safely implemented in clinical gastroenterology practice.",
  "authors": [
    {
      "affiliations": [
        "University of Bristol, Bristol, United Kingdom"
      ],
      "name": "Jason Ha"
    },
    {
      "affiliations": [
        "King’s College London, London, United Kingdom"
      ],
      "name": "Joshua Lee"
    },
    {
      "affiliations": [
        "King’s College London, London, United Kingdom"
      ],
      "name": "Victoria To"
    },
    {
      "affiliations": [
        "King’s College London, London, United Kingdom"
      ],
      "name": "Ayush Patel"
    },
    {
      "affiliations": [
        "King’s College London, London, United Kingdom"
      ],
      "name": "Vedant Bhardwaj"
    },
    {
      "affiliations": [
        "Guy’s and St Thomas’ NHS Foundation Trust, London, United Kingdom"
      ],
      "name": "Sebastian Zeki"
    }
  ],
  "title": "P445 Accuracy and readability of large language models for gastrointestinal procedures: a systematic review",
  "uid": "4a987b6a-f814-5bd6-ba63-1228c3497b95"
}
