{
  "abstract": "Objective Meaningful assessments of how large language models (LLMs) incorporate clinical guidelines require large-scale testing over many queries. Here, we evaluate the prevalence of clinical guideline omissions and hallucinations in a large sample of diagnostic LLM outputs.Methods We used simulated case vignettes and zero-shot prompting to generate diagnostic outputs and rationales from GPT-4.1 and DeepSeek-V3. English case vignettes were created for hypercholesterolaemia and type-2 diabetes mellitus. Each vignette contained identical medical information, while sociodemographic characteristics varied in terms of sex, ethnicity and location. We calculated the prevalence of existing and hallucinated clinical guidelines in LLM outputs across disease, LLM and sociodemographic characteristics.Results We analysed a total of 12 197 LLM outputs, which quantifies three hazard areas: omissions (up to 97% for DeepSeek-V3 and 46% for GPT-4.1), hallucinations (up to 9%) and inconsistencies (guideline citation rate ranging from 0% to 78.39% across sociodemographic vignettes). Omission and hallucination rates were generally similar across vignettes with different sex or ethnicity data, yet were particularly sensitive to patient location.Discussion This study highlights significant variability in clinical guideline prediction across two different diseases, three different sociodemographic variables and two LLMs, even when the LLMs were instructed by identical prompts, establishing clinical guideline prediction in LLM outputs as a stochastic event.Conclusion The stochastic nature of LLMs creates a unique challenge for evidence generation and clinical deployment. Being able to measure and capture this stochasticity within high-quality research designs will be a prerequisite to advancing the responsible deployment of LLMs in healthcare.",
  "authors": [
    {
      "affiliations": [
        "AI and Digital Unit, LSE Health, The London School of Economics and Political Science, London, UK",
        "Centre for Health and Healthcare, World Economic Forum, Cologny, Switzerland"
      ],
      "name": "Robin van Kessel"
    },
    {
      "affiliations": [
        "AI and Digital Unit, LSE Health, The London School of Economics and Political Science, London, UK",
        "Centre for Primary Care and Health Services Research, The University of Manchester, Manchester, UK"
      ],
      "name": "Michael Anderson"
    },
    {
      "affiliations": [
        "Centre for Primary Care and Health Services Research, The University of Manchester, Manchester, UK"
      ],
      "name": "Brian McMillan"
    },
    {
      "affiliations": [
        "Department of Family Medicine, Mayo Clinic, Rochester, Minnesota, USA"
      ],
      "name": "Marc R Matthews"
    },
    {
      "affiliations": [
        "Faculty of Health, School of Medicine, Witten/Herdecke University, Witten, Germany"
      ],
      "name": "Paul Rust"
    },
    {
      "affiliations": [
        "AI and Digital Unit, LSE Health, The London School of Economics and Political Science, London, UK"
      ],
      "name": "Pauline Pearcy"
    },
    {
      "affiliations": [
        "Division of Prevention and Wellness & Center for Cardiovascular Computational & Precision Health, DeBakey Heart & Vascular Center, Houston Methodist, Houston, Texas, USA",
        "Houston Methodist-Rice Digital Health Institute, Houston, Texas, USA"
      ],
      "name": "Khurram Nasir"
    },
    {
      "affiliations": [
        "AI and Digital Unit, LSE Health, The London School of Economics and Political Science, London, UK"
      ],
      "name": "Elias Mossialos"
    }
  ],
  "title": "Omission and hallucination prevalence of clinical guidelines in diagnostic large language model outputs",
  "uid": "649e0094-ccd6-54af-9f34-97d8022574fe"
}
