{
  "abstract": "Background Inflammatory bowel disease (IBD) patients face fertility-related concerns, yet reliable information is scarce. Large language models (LLMs) like ChatGPT and DeepSeek show promise for patient education, but their performance on IBD fertility queries is unevaluated.Methods 16 IBD fertility-related questions were posed to ChatGPT-4o and DeepSeek-R1 twice at a 2-week interval. The two IBD professionals assessed the accuracy and reproducibility of the responses. Accuracy was scored using a four-point scale (completely correct, 4 points; correct but insufficient, 3 points; mix of correct and incorrect/outdated information, 2 points; completely incorrect, 1 point). Questions with divergent ratings were resolved by a third senior specialist. Cohen’s kappa coefficient (κ) was calculated to evaluate the inter-rater consistency of the two reviewers.Results A flow diagram detailing study inclusion was shown in Figure 1 ( IDDF2026-ABS-0419 Figure 1. Flow diagram detailing study inclusion). Both LLMs achieved high accuracy, with 100% of responses rated as ‘completely correct’ or ‘correct but insufficient’, and no significant difference in the overall accuracy score between the two models (3.34±0.49 vs. 3.28±0.46; P=0.536), shown in Table 1 (IDDF2026-ABS-0419 Table 1) and Table 2 (IDDF2026-ABS-0419 Table 2). Figure 2 (IDDF2026-ABS-0419 Figure 2. Grade of responses by ChatGPT-4o language model) and Figure 3 (IDDF2026-ABS-0419 Figure 3. Grade of responses by DeepSeek-R1 language model) illustrate the response capability percentages for both models. DeepSeek-R1 performed better in the areas of genetics, pregnancy medication use and delivery modes, while ChatGPT-4o had superior performance in fertility capacity. Both showed high reproducibility with consistent core content across repeated queries. DeepSeek-R1 provided more detailed updates on four questions in the second response, and ChatGPT-4o’s details varied by topic (IDDF2026-ABS-0419 Table 1). Cohen’s κ=0.714 indicated good inter-rater reliability.Conclusions ChatGPT and DeepSeek are reliable auxiliary tools for reproductive-age IBD patients seeking fertility information. Both language models demonstrate robust and comparable performance, particularly DeepSeek with its integration of up-to-date evidence, showing promise in enhancing patient education quality, alleviating anxiety, and improving treatment adherence.Abstract IDDF2026-ABS-0419 Figure 1Abstract IDDF2026-ABS-0419 Table 1Accuracy and repeatability of the responses from the ChatGPT-4o and DeepSeek-R1 language modelNumberQuestionReproducibilityCG-FirstCG-SecondDS-FirstDS-SecondInheritance1Is IBD genetically predisposed?Similar33432What is the risk of passing on IBD to a child?Similar3444Fertility1Does IBD influence men’s fertility?Similar33332Is fertility reduced in a woman suffering from IBD?Similar34333Do patients with IBD require specialized assisted reproductive technologies support?Similar4333Disease activity and preparation for pregnancy1What level of disease activity should IBD patients achieve when preparing for pregnancy?Similar44342Does it recommend postponing pregnancy during periods of disease activity?Similar44443How to monitor and manage disease activity during pregnancy?Similar3334Medication use during pregnancy1Which medications are considered safe and recommended during preconception and pregnancy?Similar33342Which medications should be avoided during preconception and pregnancy? What are the potential risks associated with these medications?Similar3333Mode of delivery1How should the delivery mode be chosen for patients with IBDSimilar33332Is vaginal delivery associated with an increased risk of perianal complications? Under what conditions is a caesarean section recommended?Similar3444Pregnancy outcome1Does disease activity increase the risk of adverse pregnancy outcomes (such as miscarriage, preterm birth, or low birth weight)?Similar34432Is there an increased risk of congenital defects in infants born to parents with IBD?Similar3333Medication use during breastfeeding1Which medications are considered safe and recommended during breastfeeding?Similar33332Which medications should not be used during breastfeeding? What are the potential risks associated with these medications?Similar3333CG-First: ChatGPT-4o first time, CG-Second: ChatGPT-4o second time, DS-First: DeepSeek-R1 first time, DS-Second: DeepSeek-R1 second timeAbstract IDDF2026-ABS-0419 Table 2Coring situation and difference in grading between ChatGPT-4o and DeepSeek-R1DomainChatGPT-4oDeepSeek-R1P valueAll (±SD)3.28 ± 0.463.34 ± 1.480.597Inheritance3.25 ± 0.503.75 ± 0.500.207Fertility3.33 ± 0.523.00 ± 0.000.174Disease Activity and Preparation for Pregnancy3.67 ± 0.523.67 ± 0.521.00Medication Use During Pregnancy3.00 ± 0.003.25 ± 0.500.391Mode of delivery3.25 ± 0.503.50 ± 0.580.537Pregnancy Outcome3.25 ± 0.503.25 ± 0.501.00Medication Use During Breastfeeding3.00 ± 0.003.00 ± 0.001.00Abstract IDDF2026-ABS-0419 Figure 2Abstract IDDF2026-ABS-0419 Figure 3",
  "authors": [
    {
      "affiliations": [
        "Department of Gastroenterology, The Sixth Affiliated Hospital of Sun Yat-sen University, China"
      ],
      "name": "Jun Deng"
    },
    {
      "affiliations": [
        "Department of Gastroenterology, The Sixth Affiliated Hospital of Sun Yat-sen University, China"
      ],
      "name": "Tao Li"
    },
    {
      "affiliations": [
        "Department of Gastroenterology, The Sixth Affiliated Hospital of Sun Yat-sen University, China"
      ],
      "name": "Jia-Yin Yao"
    },
    {
      "affiliations": [
        "Department of Gastroenterology, The Sixth Affiliated Hospital of Sun Yat-sen University, China"
      ],
      "name": "Min Zhan"
    },
    {
      "affiliations": [
        "Department of Gastroenterology, The Sixth Affiliated Hospital of Sun Yat-sen University, China"
      ],
      "name": "Min Zhi"
    },
    {
      "affiliations": [
        "Section of medical affairs, The Third Affiliated Hospital of Guangzhou Medical University, China"
      ],
      "name": "Ming Wei"
    },
    {
      "affiliations": [
        "Department of Gastroenterology, The Sixth Affiliated Hospital of Sun Yat-sen University, China"
      ],
      "name": "Xiang Peng"
    }
  ],
  "title": "IDDF2026-ABS-0419 Evaluating the accuracy and reproducibility of large language models in responding to fertility-related queries in inflammatory bowel disease: a comparative study of chatgpt and deepseek",
  "uid": "6c404fd2-53b6-5b90-9b89-59713fc123dc"
}
