{
  "abstract": "Introduction Large language models (LLMS) are revolutionising natural language processing workflows; however, their applicability in clinical settings remains uncertain. This study investigates whether open-weight LLMs demonstrate appropriate metacognitive calibration while identifying patients with inflammatory bowel disease (IBD).Methods Secondary care electronic health records from a large UK hospital were utilised, including endoscopy reports, histopathology reports, and clinic letters, under full REC ethics approval. The records were extracted, redacted, and reviewed manually by a consultant with 15 years of experience. Open-weight LLMs (DeepSeek R1 14B, DeepSeek R1 32B, Qwen 32B, Mixtral 7B, Med 42 8B) were deployed via a modular API, with prompts configured for zero-, single-, or dual-shot in-context learning (iCL), multiple narrative orders, and two temperature settings (0.5/0.75 for ‘thinking’ models and 0.75/1.0 for ‘non-thinking’ models). The models evaluated vertices including ‘likelihood of IBD,’ ‘certainty,’ and ‘complexity’ ratings. Primary outcome measures included macro F1 score; secondary measures comprised recall, precision, specificity, Matthews correlation coefficient (MCC), negative predictive value (NPV), and Brier score. Thematic error analysis was conducted using latent Dirichlet allocation (LDA) and non-negative matrix factorisation (NMF).Results A total of 1,612 patients with chronologically linked endoscopic and histopathological records were identified. These records comprised 9,311 free-text documents. Response dropout rates decreased with model size, as shown in figure 1. The LLMs achieved overall macro-F1 scores of 0.68–0.74 for IBD detection, with DeepSeek 14B marginally surpassing the larger models. Model ‘certainty’ dipped in a grey zone of ‘likelihood of IBD’ scores between 3 and 6, but remained inappropriately high. LLM estimated case ‘complexity’ increased along with ‘likelihood of IBD’, suggesting, contrary to clinical expectations, that LLMs see clear-cut IBD cases as more complex than more ambiguous ones. These findings suggest substantial metacognitive biases present within LLMs. LDA and NMF analysis of false positives and negatives revealed that LLMs frequently misclassified quiescent or inactive IBD as absent and tended to overinterpret inflammatory symptoms as indicative of IBD, suggesting that LLMs suffer from their own form of Dunning-Kruger syndrome.Conclusions While open-weight LLMs can aid in clinical cohort identification, they exhibit substantial metacognitive shortcomings, including overconfidence, clinical naivety, temperature and prompt-related effects. Further studies should explore Methods to mitigate these downsides. In the meantime, they should be used with great caution for clinical tasks.Abstract P199 Figure 1FM response dropout rates",
  "authors": [
    {
      "affiliations": [
        "University Hospital Southampton, Southampton, United Kingdom",
        "University of Southampton, Southampton, United Kingdom"
      ],
      "name": "Matt Stammers"
    },
    {
      "affiliations": [
        "University Hospital Southampton, Southampton, United Kingdom",
        "University of Southampton, Southampton, United Kingdom"
      ],
      "name": "Markus Gwiggner"
    },
    {
      "affiliations": [
        "University of Southampton, Southampton, United Kingdom"
      ],
      "name": "Reza Nouraei"
    },
    {
      "affiliations": [
        "University of Southampton, Southampton, United Kingdom"
      ],
      "name": "Cheryl Metcalf"
    },
    {
      "affiliations": [
        "University of Southampton, Southampton, United Kingdom"
      ],
      "name": "James Batchelor"
    }
  ],
  "title": "P199 Overconfident and under-informed: dunning–kruger–like metacognitive failures in foundation models during clinical IBD cohort identification",
  "uid": "e198e8e2-75f1-59ac-9ce8-7d8439e85526"
}
