{
  "abstract": "Objective To determine the ability of large language models (LLMs) to answer open-ended healthcare questions (rather than multiple-choice questions) accurately in a low-resource setting.Methods and analysis We benchmarked five LLMs (GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma and o3) against Kenyan clinicians, using a randomly subsampled dataset of 507 vignettes (from a larger pool of 5107 clinical scenarios) spanning 12 nursing competency categories. Blinded physician panels rated responses using a 5-point Likert scale on an 11-domain rubric covering accuracy, safety, contextual appropriateness, and communication. We summarised mean scores and used Bayesian ordinal logistic regression to estimate probabilities of high-quality ratings (≥4) and to perform pairwise comparisons between LLMs and clinicians.Results Clinician mean ratings were lower than those for LLMs in 9/11 domains: 2.86 vs 4.25–4.72 (guideline alignment), 2.76 vs 4.25–4.73 (expert knowledge), 2.96 vs 4.30–4.73 (logical coherence) and 2.58 vs 4.16–4.68 (low omission of critical information). On safety-related domains, LLMs received higher ratings: minimal extent of possible harm 3.16 vs 4.29–4.68; low likelihood of harm 3.68 vs 4.54–4.81. Performance was similar for low inclusion of irrelevant content (4.28 vs 4.25–4.35) and for avoidance of demographic bias (4.86 vs 4.91–4.94). In Bayesian models, LLMs had >90% probability of ratings ≥4 in most domains, whereas clinicians exceeded 90% only for contextual relevance and demographic/socioeconomic bias. Pairwise contrasts showed broadly overlapping credible intervals among LLMs, with o3 leading numerically most domains except contextual relevance, demographic/socio-economic bias and relevance to the question. Generating all LLM responses cost US$3.86–US$8.68 per model (US$0.008–US$0.017 per vignette), compared with US$3.35 per clinician-generated vignette.Conclusions In controlled vignette-based tasks, LLMs produced responses that were more accurate, safer and more structured than clinicians, suggesting LLMs may have potential as supplementary knowledge and safety support tools. Findings support further evaluation in real patient encounters to determine effectiveness, safety and integration into clinical workflows, particularly in resource-constrained health systems.",
  "authors": [
    {
      "affiliations": [
        "KEMRI-Wellcome Trust Research Programme, Nairobi, Kenya",
        "Keprecon, Nairobi, Kenya"
      ],
      "name": "Paul Mwaniki"
    },
    {
      "affiliations": [
        "PATH, Nairobi, Kenya"
      ],
      "name": "Wilkister Musau"
    },
    {
      "affiliations": [
        "KEMRI-Wellcome Trust Research Programme, Nairobi, Kenya",
        "Keprecon, Nairobi, Kenya"
      ],
      "name": "Lynda Isaaka"
    },
    {
      "affiliations": [
        "KEMRI-Wellcome Trust Research Programme, Nairobi, Kenya",
        "Keprecon, Nairobi, Kenya"
      ],
      "name": "Conrad Wanyama"
    },
    {
      "affiliations": [
        "University of Birmingham, Birmingham, UK"
      ],
      "name": "Vaishnavi Menon"
    },
    {
      "affiliations": [
        "University of Birmingham, Birmingham, UK"
      ],
      "name": "Alastair K Denniston"
    },
    {
      "affiliations": [
        "University of Birmingham, Birmingham, UK"
      ],
      "name": "Xiaoxuan Liu"
    },
    {
      "affiliations": [
        "PATH, Geneva, Switzerland"
      ],
      "name": "Mira Emmanuel-Fabula"
    },
    {
      "affiliations": [
        "PATH, London, UK"
      ],
      "name": "Gwydion Williams"
    },
    {
      "affiliations": [
        "University of Birmingham, Birmingham, UK"
      ],
      "name": "Bilal Akhter Mateen"
    },
    {
      "affiliations": [
        "KEMRI-Wellcome Trust Research Programme, Nairobi, Kenya",
        "Keprecon, Nairobi, Kenya",
        "LSHTM, London, UK"
      ],
      "name": "Ambrose Agweyu"
    }
  ],
  "title": "Benchmarking large language models and clinicians using locally generated primary healthcare vignettes in Kenya",
  "uid": "08f90655-1b92-51ed-b1a0-cf5c3bbe0d94"
}
