{
  "abstract": "Background The emergence of large language models (LLMs) represents a paradigm shift in the pharmaceutical industry, fundamentally transforming processes ranging from target identification and preclinical research to clinical trial optimization. By systematically mining both public datasets and the extensive corpus of scientific literature, LLMs provide researchers with direct, scalable access to critical knowledge, while analysis of proprietary datasets affords organizations a distinctive competitive advantage through the generation of novel, actionable insights. Within large organizations, successful biomarker development hinges on two priorities: maximizing the utility and interpretability of proprietary data, and promoting fluid knowledge exchange across multidisciplinary teams. LLMs directly support these objectives by bridging expertise from disparate domains, distilling complex scientific concepts into accessible language, and facilitating cross-functional collaboration. This integration fosters the emergence of the ‘augmented scientist’—a researcher empowered to generate, test, and refine hypotheses rapidly by leveraging cross-disciplinary data and insights.Methods AstraZeneca’s Biomarker Navigator program embodies this principle by providing three foundational capabilities: an automated analytical framework for biomarker discovery; a rigorously curated, standardized repository of biomarker insights; and an intuitive user interface to ensure rapid and effortless access to knowledge. The program is now being expanded to benefit the wider scientific community through the integration of high-value public datasets, initially targeting indications such as renal cell carcinoma (RCC) and melanoma. This expansion encompasses whole exome sequencing (WES), copy number alteration (CNA) data, RNA sequencing (RNA-seq), and richly annotated clinical metadata.Results The enhanced platform empowers researchers to traverse seamlessly from broad, landscape-level analyses across multiple datasets to in-depth explorations of individual biomarkers, ensuring no loss of critical information during the workflow.Conclusions By coupling a user-focused interface with advanced, LLM-driven data synthesis, AstraZeneca is demonstrating the transformative impact of artificial intelligence on biomarker discovery, and is providing a robust, framework to accelerate scientific innovation and decision-making.",
  "authors": [
    {
      "affiliations": [
        "AstraZeneca, Bethesda, MD, USA"
      ],
      "name": "Mishka Gidwani"
    },
    {
      "affiliations": [
        "AstraZeneca, Gaithersburg, MD, USA"
      ],
      "name": "Iaroslav Korkodinov"
    },
    {
      "affiliations": [
        "AstraZeneca, Gaithersburg, MD, USA"
      ],
      "name": "Georgiy Ponomarev"
    },
    {
      "affiliations": [
        "AstraZeneca, Cambridge, MA, USA"
      ],
      "name": "Gerald Sun"
    },
    {
      "affiliations": [
        "AstraZeneca, New York, NY, USA"
      ],
      "name": "Damian Bikiel"
    },
    {
      "affiliations": [
        "AstraZeneca, Waltham, MA, USA"
      ],
      "name": "Etai Jacob"
    },
    {
      "affiliations": [
        "AstraZeneca, San Francisco, CA, USA"
      ],
      "name": "Ioannis Kagiampakis"
    }
  ],
  "title": "1093 Empowering biomarker discovery and knowledge generation with large language models: the biomarker navigator program",
  "uid": "9e823b21-172f-582f-8171-cfb5546f5d99"
}
