Independent benchmark analysis / ICML 2025 / arXiv v3

MedXpertQA Text and MM

Expert exam difficulty needs an equally careful comparison.

MedXpertQA pairs a text examination benchmark with a multimodal counterpart. Its questions are selected and augmented to stress medical knowledge and reasoning, with different answer-choice counts across the two tracks. We analyze why that construction changes score interpretation, distinguish the test sets from their small development sets, and preserve the paper’s sampled-versus-full evaluation caveat. Our result panels intentionally use full-set rows from the historical study. The site contributes a comparison framework and original interpretation of the benchmark, while the dataset, task definitions and measured results remain credited to its creators.

01 / What is being tested?

The task, before the score.

input
Medical examination question; the MM track also supplies associated images and clinical context
output
One answer choice: ten options for Text, five for MM
unit
One examination question
setting
Original zero-shot chain-of-thought evaluation; greedy decoding where supported

Data origin. Public licensing and specialty-exam questions, filtered for difficulty, rephrased and augmented with expert review. [1][2][3]

Total release
4,460 questions

Text and MM together, including development questions.

§3.1 [1]
Text test
2,450 questions

Five additional development questions are separate.

§3.1; Table 3 [1][3]
MM test
2,000 questions

Five additional development questions are separate.

§3.1; Table 2 [1][3]
Medical specialties
17

The source defines coverage, not equal numbers per specialty.

§3.1 [1]
Answer options
10 Text / 5 MM

Different uniform-guess reference rates follow arithmetically.

Question and Option Augmentation [1]
  1. 01

    Select Text or MM

    Keep their test identifiers and development examples separate.

    [3]
  2. 02

    Apply the source prompt protocol

    Zero-shot CoT; model-specific exceptions are documented.

    [1]
  3. 03

    Extract the final option

    Answer extraction is part of the evaluation implementation.

    [2]
  4. 04

    Report track and sample coverage

    Give the actual tested population, including any sampled subset.

    [1]

Dataset anatomy

Release accounting

Text test

Excludes five development items.

2,450 questions[1]
MM test

Excludes five development items.

2,000 questions[1]
Text development

Few-shot development examples.

5 questions[1]
MM development

Few-shot development examples.

5 questions[1]

All four partitions sum to the 4,460-question release. The image total is omitted because the paper’s narrative and comparison table disagree. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Multiple-choice accuracy

Higher is better

Extract the selected option and compare it with the reference; inspect Text and MM separately.

Scoring definition

Accuracy = correct options / evaluated questions

Original o1 and o3-mini rows use sampled subsets; do not silently compare them as full-test measurements. [1]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Published full Text results

ICML 2025 / arXiv v3, Text test, zero-shot CoT; model-specific protocol exceptions apply.

Accuracy · %
050100
Reported
DeepSeek-R1Text Avg; full evaluation.
37.76%
GPT-4ogpt-4o-2024-11-20; Text Avg.
30.37%
DeepSeek-V3Text Avg; full evaluation.
24.16%

Selected historical paper-reported rows. No current leaderboard claim and no uncertainty supplied in the selected table.

Source: Tables 4–5 [1]

Paper-reported results / selected rows

Published full MM results

Same 2025 paper; multimodal test under its original prompt and answer-extraction protocol.

Accuracy · %
050100
Reported
GPT-4ogpt-4o-2024-11-20; MM Avg.
42.80%
Gemini-2.0-FlashMM Avg.
37.20%
Claude-3.5-SonnetMM Avg.
33.20%

Separate questions and answer-choice count from Text. A higher MM score does not show that images improved the same cases.

Source: Table 4 [1]

04 / Our original analysis

What follows from the design?

01

Different tracks have different chance references

Published evidence

Text has ten choices and MM has five. [1]

Our interpretation

Our inference: uniform guessing yields 10% and 20% respectively; raw cross-track score gaps cannot isolate a visual benefit.

02

Difficulty selection defines the target population

Published evidence

The benchmark filters and augments examination questions. [1]

Our interpretation

Our inference: failure rates here describe a deliberately challenging exam population, not the prevalence of errors in routine clinical work.

03

Coverage matters before ranking

Published evidence

The paper evaluates some expensive models on sampled subsets. [1]

Our interpretation

Our inference: retain an explicit full-versus-sampled field even when all scores share the same benchmark name.

04

Rephrasing is a mitigation, not a certificate

Published evidence

Public questions are augmented to reduce leakage risk. [1]

Our interpretation

Our inference: changed wording does not by itself prove that the underlying question or solution was absent from training.

05 / Scope of the evidence

Where this benchmark stops.

Examination setting

Selecting a supplied option does not test open-ended information gathering or treatment execution. [1]

Public source exposure

Public question provenance and augmentation leave residual contamination uncertainty. [1]

Human reference scope

The paper’s pre-licensed expert row is not a universal practicing-clinician baseline. [1]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public official repository and dataset
License
MIT declared by the official dataset card
Conditions
Preserve attribution and distinguish repository terms from any rights in source exam material.
[3]

Evidence trail

Read the originals.

  1. MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding ↗

    Zuo et al. / Tsinghua University and Shanghai AI Laboratory. Versioned construction, options, zero-shot evaluation and selected result tables.

  2. MedXpertQA official repository ↗

    TsinghuaC3I. Official release links and evaluation implementation.

  3. MedXpertQA official dataset card ↗

    TsinghuaC3I. Separate Text/MM dev and test configurations; MIT license declaration.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗