{
  "abstract": "As the capabilities of large language models (LLMs) continue to improve, AI chatbots may develop strong persuasive capabilities. As such, there is growing interest in the use of interactive, LLM-powered tools for smoking cessation; examples include the World Health Organization‘s chatbot Sarah and BeFreeGPT and BasicGPT, which were developed recently by a group at George Washington University. As with other applications of generative AI to clinical and public health communication tasks, there is a need to evaluate both the accuracy of information produced and whether it is tailored appropriately to the intended audience.Previous work has established that AI chatbots can produce accurate information about smoking cessation (Abrams et al. 2025). In this paper, we apply a suite of computational methods, as well as user testing with online participants, to investigate accessibility and tailoring of information. For 12 common internet queries about ‘how to quit smoking,’ we generated three responses per query using Sarah, BeFreeGPT, and BasicGPT. We then evaluated these responses using a multimodal approach encompassing i) natural language processing metrics for syntactic complexity, jargon, and lexical cohesion, ii) structured semi-quantitative assessments such as the Patient Education Materials Assessment Tool and the Clear Communication Index, and iii) direct user testing. For the user studies, we recruited participants who reported smoking through Prolific, with quota sampling to achieve samples representative of the U.S. population for age, race, ethnicity, and education. Participants, who were randomly assigned either a human-written control passage or a passage generated by BeFreeGPT, answered reading comprehension questions about the material provided and rated its persuasiveness using a Likert scale.For all three sets of passages, we observed substantial heterogeneity across our different modes of evaluation. Of particular interest, we found that our additional computational metrics often yielded divergent predictions from length-based readability formulas, which remain standard in the field. AI-generated passages, however, were rated as significantly more persuasive in the user study, suggesting that AI-assistant approaches to smoking cessation warrant further investigation.",
  "authors": [
    {
      "affiliations": [
        "University of Macau, Macau, China"
      ],
      "name": "JP Dexter"
    }
  ],
  "title": "M4 Multimodal evaluation of AI-assisted approaches to smoking cessation",
  "uid": "37806644-7425-5d84-baf6-d16c7aaa6686"
}
