The emergence of "AI Scientist" systems capable of autonomous research promises to accelerate discovery, but it raises a critical question: how do we evaluate the quality of AI-generated papers? A new study published on ArXiv (cs.AI) proposes a rigorous benchmarking protocol using an automated peer-review system powered by frontier large language models (LLMs).

The Benchmarking Protocol

Researchers evaluated four leading AI Scientist frameworks: Sakana AI (v1 & v2), CycleResearcher, and Data-to-Paper. The benchmark utilized 15 research proposals from the commercial AI company FARS, generating 60 papers that were evaluated alongside 15 FARS benchmark papers. The assessment focused on four core dimensions:

  • Originality
  • Scientific rigor
  • Clarity
  • Significance

Results and the GPT-5.4 Divergence

The findings reveal that FARS benchmark papers significantly outperform competing frameworks, achieving mean scores of 2.14–2.47 on a 1–5 scale, compared to 1.00–1.87 for other systems. Notably, FARS scores were more than double those of the next-best systems in evaluations by Gemini and Claude. A key technical observation was the correlation between reviewers. While Gemini and Claude showed strong agreement (ρ = 0.907), GPT-5.4 exhibited much weaker agreement (ρ ≈ 0.32), suggesting it employs different criteria for assessing scientific quality. Despite this, the study concludes that multi-model LLM evaluation provides a scalable and consistent framework for assessing autonomous research.