{
  "abstract": "Introduction Artificial intelligence is increasingly proposed for clinical decision support in acute ischemic stroke (AIS). The 2026 AIS Guidelines introduced important framework changes — including ASPECTS-based endovascular thrombectomy (EVT) eligibility independent of perfusion imaging. We compared a purpose-built deterministic clinical decision support system (MedSync AI) against leading probabilistic large language models (LLMs) to evaluate clinical reliability and guideline concordance for acute stroke treatment.Materials and Methods Twenty-one standardized clinical vignettes were constructed spanning the full IVT and EVT decision range of the 2026 guidelines — including straightforward cases, conditional cases, not-recommended cases, guideline gaps, and missing data scenarios. Five systems were evaluated: MedSync AI (deterministic rule-based), Claude with guidelines, ChatGPT-5.3 with guidelines, Claude alone, and ChatGPT-5.3 alone, each receiving identical point-of-care prompts.Responses were scored using a validated three-condition framework. Faithful Representation — the primary metric — required all three conditions simultaneously: (1) Correct Direction: recommendation points the right way; (2) Correct Reasoning: logic maps to the applicable 2026 guideline recommendation and its specific eligibility criteria, not to superseded frameworks or non-mandatory requirements; and (3) Correct Evidence Level: language conveys the appropriate Class of Recommendation through clinical certainty and tone — a COR 1 must be definitive and unqualified, a COR 3 clearly negative. A correct answer reached through flawed reasoning scores as unfaithful. Correct Clinical Correlation (CCC) — our operationalization of reasoning faithfulness — combined conditions 1 and 2. The Direction−CCC Gap quantified responses with correct direction produced through flawed reasoning — a prospective patient safety risk. Formal Fidelity required explicit COR, LOE, and recommendation citation.Results MedSync AI achieved 100% across all metrics with zero wrong-direction errors and a Direction−CCC Gap of zero — every correct answer produced by correct reasoning, auditable and reproducible. LLM performance fell substantially short of standards required for a reliable healthcare AI clinical decision tool.Even when provided the 2026 guidelines directly, Claude reached only 84.2% Faithful Representation and ChatGPT-5.3 only 55.0% — failing nearly half of all scenarios under the full three-condition standard. Without guidelines, performance collapsed: Claude 40.0% and ChatGPT-5.3 24.0%, with 5 and 8 wrong-direction errors respectively — recommending against guideline-indicated thrombectomy in patients who clearly qualified. The most dangerous error occurred in both standalone conditions: ChatGPT classified DOAC anticoagulation as an absolute contraindication to IVT with COR 3 Harm designation, directly contradicting the 2026 guideline’s explicit classification as a relative contraindication requiring individualized assessment. At clinical safety rates of 76% and 71%, nearly 1 in 4 standalone LLM responses contained an element capable of causing patient harm. No LLM achieved Formal Fidelity under any condition.Conclusion General-purpose LLMs — even when provided direct guideline access — did not meet the reliability standards required for a clinical AI decision tool in acute stroke care. Only a purpose-built deterministic system achieved the guideline concordance, reasoning consistency, and auditability that point-of-care stroke decisions demand. As AI enters clinical workflows, architecture matters: a system that produces the right answer through the wrong reasoning is not a tool physicians can trust.Disclosures M. Stiefel: None. D. Stiefel: None.",
  "authors": [
    {
      "affiliations": [
        "MedSync AI, Charlotte, NC"
      ],
      "name": "M Stiefel"
    },
    {
      "affiliations": [
        "MedSync AI, Charlotte, NC"
      ],
      "name": "D Stiefel"
    }
  ],
  "title": "O-047 Development of a scalable AI-driven clinical decision tool for acute stroke: a comparison of deterministic rule-based and probabilistic large language models",
  "uid": "fb448850-6086-57c4-8352-99b8265a7f11"
}
