New Benchmark PHYBench Challenges Large Language Models’ Reasoning Skills

Neural Edition

Artificial Intelligence

New Benchmark PHYBench Challenges Large Language Models' Reasoning Skills

PHYBench introduces 500 original physics problems to assess the reasoning capabilities of large language models more effectively.

Artificial IntelligenceWorking knowledge2 min read

New Benchmark PHYBench Challenges Large Language Models' Reasoning Skills

Featured image: Desk-office-workspace-coworking (23699033283).jpg by www.Pixel.la Free Stock Photos, licensed under CC0.

A group of researchers has launched PHYBench, a new benchmark aimed at rigorously evaluating the reasoning capabilities of large language models (LLMs) through 500 original physics problems.

PHYBench: Gemini 2.5 Pro vs Human ExpertsGemini 2.5 Pro36.9% accuracy#1Human Experts61.9% accuracy#2
On PHYBench’s 500 original physics problems, Gemini 2.5 Pro reaches 36.9% accuracy while human experts reach 61.9%.

What Happened

The researchers introduced PHYBench as a solution to the limitations in existing benchmarks for evaluating reasoning in LLMs, noting that many current methods oversimplify tasks and suffer from data contamination.

The Backstory

Current benchmarks for reasoning in LLMs often fail to accurately gauge the reasoning abilities of models due to several key issues:

  • Task Oversimplification: Existing tests may not encompass the full range of problem difficulty.
  • Data Contamination: Many benchmarks include data that can skew evaluations.
  • Flawed Evaluation Items: Errors in existing problems can misrepresent a model’s capabilities.

What are we talking about?

  • LLMs: Models that generate human-like text based on input.
  • Data Contamination: Inclusion of biased or error-prone data that affects evaluation.
  • Reasoning Skills: The ability to engage in complex thought processes and problem-solving.

How It Works

  1. Development of Original Problems: Researchers created 500 unique physics problems covering varying levels of difficulty.
  2. Systematic Curation: A pipeline was established to ensure the elimination of flawed items and contamination.
  3. Evaluation Methodology: The benchmark was tested against existing methods like AIME 2024 and OlympiadBench.
  4. Introduction of EED Score: The Expression Edit Distance Score facilitates more precise assessment of mathematical expressions.
  5. Multi-step and Multi-condition Reasoning: The benchmark aims to assess the robustness of reasoning across complex task requirements.
Original Problems
Systematic Curation
EED Score Introduction
Overall Evaluation

The Numbers

Evaluation results indicated that the best-performing LLM, Gemini 2.5 Pro, achieved only 36.9% accuracy on PHYBench. In contrast, human experts scored an average of 61.9%. These results highlight the gap between machine reasoning and human expertise.

What Changed

Gemini 2.5 Pro36.9% Accuracy
Human Experts61.9% Accuracy

Why it Matters

  • PHYBench provides a clearer assessment of reasoning capabilities in LLMs.
  • The findings point to significant gaps in current LLM performance compared to human reasoning.
  • It establishes a basis for further research into enhancing reasoning in LLMs.

What This Does Not Mean

The limited accuracy of LLMs on PHYBench does not imply that these models are ineffective overall. Rather, it indicates specific areas where improvements are needed in machine reasoning and problem-solving.

What Happens Next

Researchers will likely focus on refining LLM architectures to better meet the challenges presented by PHYBench. Ongoing evaluations could further guide improvements in the models’ performance in various reasoning tasks.

End-to-End Recap

  1. PHYBench introduces 500 original physics problems.
  2. It addresses limitations found in current benchmarks.
  3. Evaluation highlights considerable gaps in LLM performance.
  4. The EED Score enhances assessment accuracy significantly.
  5. Further research aims to improve LLM reasoning skills.

Learn · Try · Watch

  • learn

    Study PHYBench: Holistic Evaluation

    Explore the methods and implications of the new PHYBench for LLM evaluation.

  • try

    Test reasoning problems

    Engage with sample problems from PHYBench to understand model performance.

    About 20 minutes.

  • watch

    Monitor LLM performance trends

    Track the progress of LLMs as they adapt to new benchmarks like PHYBench.

    What matters: Look for improving accuracy and reasoning skills over time.

  • try today

    Build a five-case eval table

    Pick one prompt you reuse. Write five rows: input, expected behavior, and pass/fail. Run them once today and keep the table next to the prompt.

    About 20 minutes.

Editor’s note: Neural Edition summarizes public reporting and labels company or founder claims as such. How we report · Corrections