{
  "abstract": "Objectives To evaluate whether variability in patients’ communication style (personality), international English-accents (human and synthetic) and speech impairments affects the accuracy of a Clinical AI Scribe (CAIS) and identify where performance degrades to inform pre-deployment validation and monitoring.Methods We conducted simulated primary-care consultations using trained actors. For personality types, four scenarios were enacted, each with five patient-personality types. For accents, transcripts of consultations were used to generate combinations of seven accents across five scenarios. The CAIS produced summaries that were compared with transcripts, and errors classified as omissions, factual inaccuracies or hallucinations. For speech impairments, public recordings representing five profiles were transcribed and word-recognition accuracy calculated.Results Personality types showed no statistically significant differences in errors (all p>0.05). Extraversion had the highest total errors (median 3.5). Across accents, comparisons were non-significant for both patient and doctor voices (patients: p=0.851; doctors: p=0.980). Omissions predominated, with low rates of hallucinations and factual inaccuracies. Omissions were slightly higher for Chinese-accented and Indian-accented doctors (both medians 3.0). Conversely, speech impairments differed: cleft palate and vowel disorders were near-perfect, whereas phonological impairment markedly reduced recognition (p<0.001).Discussion Operationally, CAIS deployment should include clinician-in-the-loop verification, subgroup performance monitoring (accents, impairments) and predefined ‘switch-off’ criteria for severe phonological patterns. High-quality synthetic voices are a pragmatic proxy for accent testing when balanced corpora are unavailable.Conclusions Under controlled conditions CAIS performance was broadly stable across communication styles and most accents, but vulnerable to specific speech characteristics, particularly phonological impairment, in this single-system simulation study.",
  "authors": [
    {
      "affiliations": [
        "University of the West of England, Bristol, UK"
      ],
      "name": "Thomas C Draper"
    },
    {
      "affiliations": [
        "University of the West of England, Bristol, UK"
      ],
      "name": "Jason Leake"
    },
    {
      "affiliations": [
        "University of the West of England, Bristol, UK"
      ],
      "name": "Timothy Cox"
    },
    {
      "affiliations": [
        "University of the West of England, Bristol, UK"
      ],
      "name": "Kathryn Lamb-Riddell"
    },
    {
      "affiliations": [
        "NHS England and NHS Improvement South West, Taunton, UK"
      ],
      "name": "Benjamin E Johns"
    },
    {
      "affiliations": [
        "NHS England and NHS Improvement South West, Taunton, UK"
      ],
      "name": "John McCormick"
    },
    {
      "affiliations": [
        "NHS England and NHS Improvement South West, Taunton, UK"
      ],
      "name": "Stephen Trowell"
    },
    {
      "affiliations": [
        "University of the West of England, Bristol, UK"
      ],
      "name": "Janice Kiely"
    },
    {
      "affiliations": [
        "University of the West of England, Bristol, UK"
      ],
      "name": "Richard Luxton"
    }
  ],
  "title": "AI-generated clinical summaries: errors and susceptibility to speech and speaker variability",
  "uid": "b7ec0288-729f-56f1-955f-a26b850b78a3"
}
