{
  "abstract": "Objective Closed-source large language models (LLMs) like generative pre-trained transformer 4o (GPT-4o) have shown promise for clinical information extraction but are potentially limited by cost, data security concerns and inflexibility. Open-source models are an attractive alternative with various adaptation strategies with no consensus on best practices. This study aims to rigorously identify optimal adaptation strategies for open-source models and evaluate their performance relative to closed-source alternatives.Methods and Analysis We studied three LLM adaptation strategies: chain-of-thought prompting, few-shot prompting and fine-tuning. Our target for information extraction was the Mayo Endoscopic Subscore (MES). We applied those strategies in all combinations to six open-source models (8–70 billion parameters) using an annotated set of colonoscopy procedure reports from the University of California, San Francisco (N=608) and San Francisco General Hospital (N=217). We analysed the relationship of these strategies to several performance metrics with a mixed-effects model, accounting for the variability between centres and LLMs. GPT-4o served as a closed-source oracle and provided in-depth commentary on the cost-effectiveness of these options.Results Quantised low-rank adaptation (QLoRA) statistically improves ( p§amp;lt;0.001)) the performance of open-source LLMs by 9.1–15.7 percentage points across accuracy, precision recall and annotation eligibility accuracy. However, GPT-4o with prompt engineering outperforms the best open-source model by 4.9%–11.2%. A simple cost-effectiveness analysis suggests that GPT-4o is more affordable compared with open-source alternatives.Conclusion GPT-4o is currently the most efficient LLM for MES extraction. If unavailable, QLoRA-optimised open-source models are a competitive alternative. However, results also suggest that current instruction-following LLMs including GPT-4o do not fully follow user-provided instructions, leaving room for improvement. More work is needed to achieve consistent, near-perfect performance in clinical information extraction by LLMs.",
  "authors": [
    {
      "affiliations": [
        "Bakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, California, USA"
      ],
      "name": "Richard Paul Yim"
    },
    {
      "affiliations": [
        "Mayo Clinic Division of Gastroenterology and Hepatology, Scottsdale, Arizona, USA",
        "Icahn School of Medicine at Mount Sinai Division of Gastroenterology, New York, New York, USA"
      ],
      "name": "Anna L Silverman"
    },
    {
      "affiliations": [
        "MS in Data Science and Artificial Intelligence Program, University of San Francisco, San Francisco, California, USA"
      ],
      "name": "Shan Wang"
    },
    {
      "affiliations": [
        "Bakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, California, USA",
        "Division of Gastroenterology, Department of Medicine, University of California, San Francisco, San Francisco, California, USA"
      ],
      "name": "Vivek A Rudrapatna"
    }
  ],
  "title": "Optimising large language models for clinical information extraction: a benchmarking study in the context of ulcerative colitis research",
  "uid": "e352f8a4-c823-5dae-bfe2-c8da6cdc5877"
}
