AstraZeneca has detailed the development of "Research Assistant," an internal system powered by Large Language Models (LLMs) designed to assist scientists and clinicians in navigating complex biomedical inquiries. The system provides a chat-style interface that unifies a diverse array of data sources, streamlining research and development (R&D) workflows across the pharmaceutical giant.

Integrated Data Ecosystem

The Research Assistant goes beyond simple text generation by synthesizing evidence from various domains, including:

  • Scientific literature and knowledge graphs.
  • Chemistry data and clinical trial records.
  • Safety resources and gene expression data.
  • Internal AstraZeneca experimental systems.

The system supports two primary operational modes: a "fast mode" for direct question answering and a "multi-step mode" designed to handle more intricate research tasks that require sequential logic and deeper exploration.

Grounding and Verification

A critical component of the system's architecture is its commitment to data integrity. Responses are grounded in retrieved evidence and maintain direct links back to the original source material. This allows researchers to verify the AI's output and further explore the underlying data. The technical documentation outlines the design choices and lessons learned from deploying this agentic system at scale to support daily scientific operations.