{
  "abstract": "Objectives The rapid evolution of large language models (LLMs) and their growing application in clinical text processing have created an urgent need for reliable de-identification mechanisms. While LLMs show promise in identifying sensitive health information (SHI), their capabilities require rigorous evaluation. This study aims to conduct a comprehensive benchmarking analysis of various LLM-based, traditional rule-based and hybrid de-identification methods.Methods Our benchmark analysis used five datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016 and OpenDeID v1) from different countries. We developed three baseline and eight LLM-based models. The experimental setup encompassed nine different settings using various combinations of training and testing sets to assess model robustness and cross-dataset performance.Results In the baseline models, the approach trained on the combined corpus of all five datasets (setting 3) significantly outperformed the other settings, achieving a strict F1 micro-average score of 0.8172. Regarding LLM-based models, the supervised fine-tuning approach using the same combined configuration (setting 9) achieved the highest performance with a strict F1 score of 0.9447.Discussion The harmonisation of corpora ensured standardised data formatting and SHI management across five diverse datasets, highlighting the necessity for uniform categorisation to enhance the reliability of de-identification results.Conclusions Our findings indicate that while fine-tuned LLMs offer superior accuracy, the observed performance variability across heterogeneous electronic health record sources poses significant technical challenges. Real-world implementation must address these inconsistencies to overcome the ethical and technical hurdles associated with deploying LLMs for handling sensitive health data.",
  "authors": [
    {
      "affiliations": [
        "CGD Health Pvt Ltd, Mumbai, Maharashtra, India"
      ],
      "name": "Omkar Panchal"
    },
    {
      "affiliations": [
        "Department of Medical Research, National Taiwan University Hospital, Taipei City, Taiwan"
      ],
      "name": "Nai-Wen Chang"
    },
    {
      "affiliations": [
        "Department of Electrical Engineering, National Kaohsiung University of Science and Technology, Kaohsiung, Taiwan"
      ],
      "name": "Zi-Rui Zhao"
    },
    {
      "affiliations": [
        "CGD Health Pvt Ltd, Mumbai, Maharashtra, India"
      ],
      "name": "Divyabharathy Ramesh Nadar"
    },
    {
      "affiliations": [
        "Department of Electrical Engineering, National Kaohsiung University of Science and Technology, Kaohsiung, Taiwan",
        "National Institute of Cancer Research, Zhunan Township, Taiwan"
      ],
      "name": "Hong-Jie Dai"
    },
    {
      "affiliations": [
        "Discipline of General Practice, School of Clinical Medicine, University of New South Wales, Sydney, New South Wales, Australia",
        "SREDH Consoritum, Sydney, New South Wales, Australia"
      ],
      "name": "Jitendra Jonnagaddala"
    }
  ],
  "title": "Benchmarking large language models for de-identification of electronic health record notes",
  "uid": "bac3a505-8c53-54c4-957a-0c09b15ffb1f"
}
