A global collaboration has launched BioEVAL (BioEngineering Validation of AI and LLMs), a new framework designed to test Large Language Models (LLMs) within the bioengineering domain. Unlike previous benchmarks that primarily measure factual memory, this initiative focuses on experimental reasoning across 11 distinct subfields.

Development and Composition

The project involved 22 research institutions creating a PhD-level evaluation set of 608 items. The benchmark includes:

  • 359 multiple-choice questions, finalized after a rigorous auditing process.
  • 218 tasks centered on literature synthesis.
  • 10 multimodal challenges involving the interpretation of experimental imagery.

To maintain high standards, the team conducted centralized quality control and a blinded audit, which resulted in the exclusion of 21 items from the final MCQ count.

Key Findings

Testing encompassed various models, including cloud-based systems like ChatGPT, Gemini, and Grok, alongside models capable of running on local hardware. The highest recorded accuracy reached 90% for multiple-choice questions and 80% for the multimodal reasoning subset. However, the researchers noted that performance varied significantly depending on the specific bioengineering subfield.

BioEVAL is structured as an evolving resource, utilizing standardized protocols to incorporate future expert contributions and model assessments.