{
  "abstract": "Background Large language models (LLMs) excel in text-based medical exams, but their ability to integrate multimodal data, critical for ophthalmology, is underexplored. This study evaluates vision–language LLMs’ accuracy and reasoning in complex ophthalmic questions.Methods We assessed three multimodal LLMs (CLM-V, ChatGPT-5, MiniCPM-V 4.5) on 316 bilingual ophthalmology questions (175 English Basic and Clinical Science Course single-choice, 141 Chinese senior professional title multiple-choice questions) across cornea, uvea, glaucoma, retina and orbit. Each question paired a clinical vignette with an image. Models were tested with reasoning-enabled and reasoning-disabled prompts. Accuracy was measured against reference standards, and reasoning quality was evaluated using automated rubric scoring (accuracy, data synthesis, logic, option analysis, safety) and expert review. Four cases were analysed qualitatively.Results Reasoning-enabled prompting increased mean artificial intelligence-assisted total scores in the English dataset from 14.97 to 16.07 for CLM-V, 20.77 to 23.97 for ChatGPT-5 and 10.83 to 12.60 for MiniCPM-V 4.5; in the Chinese dataset, the corresponding scores were 9.03 to 10.27, 19.95 to 22.00 and 11.05 to 13.30, respectively. Human evaluation showed substantial inter-rater agreement (κ=0.87) and again ranked ChatGPT-5 highest. Qualitative case analyses illustrated that reasoning-enabled outputs were often more clinically interpretable, although the magnitude of benefit was model-dependent and dataset-dependent.Conclusion Multimodal LLMs demonstrate potential in ophthalmic question-answering, with reasoning-enabled prompting being associated with improved interpretability and, in most settings, numerically higher performance. However, limitations in subspecialty robustness and image interpretation necessitate rigorous reasoning evaluation for safe educational and clinical applications.",
  "authors": [
    {
      "affiliations": [
        "Eye Center of Second Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, Zhejiang, People's Republic of China",
        "Zhejiang Provincial Key Laboratory of Ophthalmology, Zhejiang Provincial Clinical Research Center for Eye Diseases, Zhejiang Provincial Engineering Institute on Eye Diseases, Hangzhou, Zhejiang, People's Republic of China"
      ],
      "name": "Houfa Yin"
    },
    {
      "affiliations": [
        "Eye Center of Second Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, Zhejiang, People's Republic of China",
        "Zhejiang Provincial Key Laboratory of Ophthalmology, Zhejiang Provincial Clinical Research Center for Eye Diseases, Zhejiang Provincial Engineering Institute on Eye Diseases, Hangzhou, Zhejiang, People's Republic of China",
        "Hangzhou Mocular Medical Technology Inc, Hangzhou, Zhejiang, People's Republic of China"
      ],
      "name": "Kaikai Zhao"
    },
    {
      "affiliations": [
        "School of Optometry, The Hong Kong Polytechnic University, Kowloon, Hong Kong",
        "Research Centre for SHARP Vision (RCSV), The Hong Kong Polytechnic University, Kowloon, Hong Kong"
      ],
      "name": "Danli Shi"
    },
    {
      "affiliations": [
        "Department of Ophthalmology, University of Warmia and Mazury, Olsztyn, Poland",
        "Institute for Research in Ophthalmology, Foundation for Ophthalmology Development, Poznan, Poland"
      ],
      "name": "Andrzej Grzybowski"
    },
    {
      "affiliations": [
        "Eye Center of Second Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, Zhejiang, People's Republic of China",
        "Zhejiang Provincial Key Laboratory of Ophthalmology, Zhejiang Provincial Clinical Research Center for Eye Diseases, Zhejiang Provincial Engineering Institute on Eye Diseases, Hangzhou, Zhejiang, People's Republic of China"
      ],
      "name": "Kai Jin"
    }
  ],
  "title": "Evaluating reasoning in multimodal large language models for ophthalmology: a bilingual benchmark study using clinical vignettes and imaging",
  "uid": "ad61f990-909c-539e-a794-b5360e1fe0c4"
}
