{
  "abstract": "Introduction Decision support from AI is increasing, however, approvals depend on clinicians retaining responsibility. To expand access to echocardiography, non-specialists with AI support is promoted. However, are clinicians able to safely detect errors and correct the AI when incorrect guidance is given? In a prospective crossover study of reporting accuracy, we test accuracy when the decision support is correct, incorrect, and without support.Methods Specialist and non-specialist clinicians with a range of experience, trained in point-of-care echocardiography, were enrolled to report 30 focused echocardiograms using a modified BSE Level 1 reporting form.They were randomised to measure and report three sets of ten carefully curated scans a) without AI support; b) with AI support without errors, and, c) with AI support with errors.The structured reports consisted of 5 domains: LV systolic function, LV size, wall thickness, aortic root diameter, and RV function from TAPSE. The erroneous support was selected to result in a highly clinically significant false positive or false negative.Each participant had access to training videos, an optional training session via Zoom, and assistance whenever help was needed with the reporting platform. The level 1 reports were graded and between groups differences derived from mixed modelling. This study received ethical approval (ICREC 6302298).Results 30 participants, with experience ranging from 9 months to 3 years, have been recruited, of which 13 provided data.For LV function and size, participants failed to recognise incorrect AI values, resulting in significantly reduced reporting accuracy as compared with correct AI (LV function 40.6% vs 90%, p=0.04; LV size 17% vs 95.5%, p=0.03), with more modest, non-significant difference between fully manual measurements and correct AI (75% vs 90%, p=0.53; 56.7% vs 95.5%, p=0.19). For simple linear measurements, participants usually detected the incorrect AI results and hence there was little difference between reports derived from incorrect AI vs correct AI assistance (LV wall thickness 39% vs 58%, p=0.75; aortic root 76% vs 90%, p=0.81, figure 2).Conclusion Participants, for complex measurements such as LV function and size, failed to spot when AI made mistakes, resulting in reduced reporting accuracy. Understanding human and AI interactions is essential for maintaining reporting accuracy and quality.Abstract 536 Figure 1AI often struggles in patients with significant left ventricular hypertrophy. In the example above, the papillary muscle in the lateral wall is erroneously included in the endocardial trace, and the apex follows the edge of the trabeculation and not the true apexAbstract 536 Figure 2There is little difference between corrupted and accurate AI for simple linear measurements like wall thickness. For more complex measurements like LV function, participants score higher with AI support than reporting manually. Performance is worst with corrupted AI",
  "authors": [
    {
      "affiliations": [
        "Imperial College, London, United Kingdom"
      ],
      "name": "Catherine Stowell"
    },
    {
      "affiliations": [
        "Imperial College, London, United Kingdom"
      ],
      "name": "Jevgeni Jevsikov"
    },
    {
      "affiliations": [
        "Imperial College, London, United Kingdom"
      ],
      "name": "Matthew Shun-Shin"
    },
    {
      "affiliations": [
        "Imperial College, London, United Kingdom"
      ],
      "name": "Darrel Francis"
    }
  ],
  "title": "536 Can AI assist recently trained users of echocardiography in interpretation and reporting?",
  "uid": "92850a75-4be9-58cf-bd72-32770474ffc1"
}
