{
  "abstract": "Objectives This study evaluates the potential of large language models (LLMs) in clinical decision-making for atypical presentations of systemic lupus erythematosus (SLE). It assesses their ability to assist with diagnosis, further investigations, and treatment planning. Diagnostic and treatment recommendations from LLMs were compared with those of treating physicians to evaluate performance and reliability. The models used were Claude 3 Opus, GPT-4o Mini, and Gemini 2.0 Flash.Methods A PubMed search identified case reports published after April 2024 to ensure exclusion from model training data. Of 95 studies screened, 24 cases from 23 studies were included, all involving atypical SLE presentations not meeting classification criteria. Exclusion criteria were prior SLE diagnosis, insufficient diagnostic data, or unavailable full texts.Clinical data—symptoms, laboratory and imaging findings, and medical history—were standardized and entered into LLMs for diagnostic, investigative, and therapeutic suggestions. Models operated offline, using internal knowledge. Antinuclear antibody testing was analyzed as an indicator of SLE suspicion. Each model’s ability to identify SLE and its ranking in differential diagnoses were assessed. Initial and long-term treatment recommendations were compared with those of treating physicians.Results All three LLMs proposed SLE as a differential in every case. Gemini most often ranked SLE first, while Claude requested ANA more frequently; neither difference was significant. GPT prioritized ANA significantly earlier (p=0.011). Initial treatment recommendations were similar across models, favoring corticosteroids as first-line therapy. Long-term plans differed: Claude showed the greatest alignment with treating physicians, whereas GPT suggested guideline-consistent alternatives such as mycophenolate mofetil (MMF) or hydroxychloroquine (HCQ). In one case, all models recommended HCQ despite a contraindication (optic nerve atrophy), underscoring safety limitations. GPT and Gemini more often retrieved rare differentials (e.g., Aicardi-Goutières, Evans syndromes), while Claude commonly assigned direct SLE diagnoses with fewer alternatives. The absence of a non-SLE control limited false-positive and specificity assessment and raises concern for overdiagnosis, particularly given GPT’s early ANA prioritization.Abstract PO:06:152 Table 1Performance comparison of LLMs in diagnostic and treatment decisions for atypical SLE casesConclusions LLMs effectively prioritized ANA testing and supported early SLE consideration in diagnostically complex cases. They demonstrated strong diagnostic and initial treatment performance but variability in long-term management and safety awareness. Further research is needed to optimize their clinical applicability and reduce potential overdiagnosis.",
  "authors": [
    {
      "affiliations": [
        "Istanbul University-Cerrahpasa, Istanbul, Turkey"
      ],
      "name": "Beste Acar"
    },
    {
      "affiliations": [
        "Istanbul University-Cerrahpasa, Istanbul, Turkey"
      ],
      "name": "Berkay Aktas"
    },
    {
      "affiliations": [
        "Istanbul University-Cerrahpasa, Istanbul, Turkey"
      ],
      "name": "Oguzhan Omer Kizilkaya"
    },
    {
      "affiliations": [
        "Istanbul University-Cerrahpasa, Cerrahpasa Faculty of Medicine, Department of Dermatology, Istanbul, Turkey"
      ],
      "name": "Zekayi Kutlubay"
    },
    {
      "affiliations": [
        "Istanbul University-Cerrahpasa, Department of Internal Medicine, Division of Rheumatology, Istanbul, Turkey"
      ],
      "name": "Serdal Ugurlu"
    }
  ],
  "title": "PO:06:152 Can large language models support clinical decision-making in atypical SLE? A comparative analysis",
  "uid": "b239fd03-ff9f-57d0-a771-f886c4c50269"
}
