Different tracks have different chance references
Text has ten choices and MM has five. [1]
Our inference: uniform guessing yields 10% and 20% respectively; raw cross-track score gaps cannot isolate a visual benefit.
Independent benchmark analysis / ICML 2025 / arXiv v3
MedXpertQA pairs a text examination benchmark with a multimodal counterpart. Its questions are selected and augmented to stress medical knowledge and reasoning, with different answer-choice counts across the two tracks. We analyze why that construction changes score interpretation, distinguish the test sets from their small development sets, and preserve the paper’s sampled-versus-full evaluation caveat. Our result panels intentionally use full-set rows from the historical study. The site contributes a comparison framework and original interpretation of the benchmark, while the dataset, task definitions and measured results remain credited to its creators.
01 / What is being tested?
Data origin. Public licensing and specialty-exam questions, filtered for difficulty, rephrased and augmented with expert review. [1][2][3]
Text and MM together, including development questions.
§3.1 [1]The source defines coverage, not equal numbers per specialty.
§3.1 [1]Different uniform-guess reference rates follow arithmetically.
Question and Option Augmentation [1]Keep their test identifiers and development examples separate.
[3]Zero-shot CoT; model-specific exceptions are documented.
[1]Answer extraction is part of the evaluation implementation.
[2]Give the actual tested population, including any sampled subset.
[1]Dataset anatomy
Excludes five development items.
Excludes five development items.
Few-shot development examples.
Few-shot development examples.
All four partitions sum to the 4,460-question release. The image total is omitted because the paper’s narrative and comparison table disagree. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better
Extract the selected option and compare it with the reference; inspect Text and MM separately.
Accuracy = correct options / evaluated questions
Original o1 and o3-mini rows use sampled subsets; do not silently compare them as full-test measurements. [1]
03 / Measured evidence
Paper-reported results / selected rows
ICML 2025 / arXiv v3, Text test, zero-shot CoT; model-specific protocol exceptions apply.
Selected historical paper-reported rows. No current leaderboard claim and no uncertainty supplied in the selected table.
Source: Tables 4–5 [1]
Paper-reported results / selected rows
Same 2025 paper; multimodal test under its original prompt and answer-extraction protocol.
Separate questions and answer-choice count from Text. A higher MM score does not show that images improved the same cases.
Source: Table 4 [1]
04 / Our original analysis
Text has ten choices and MM has five. [1]
Our inference: uniform guessing yields 10% and 20% respectively; raw cross-track score gaps cannot isolate a visual benefit.
The benchmark filters and augments examination questions. [1]
Our inference: failure rates here describe a deliberately challenging exam population, not the prevalence of errors in routine clinical work.
The paper evaluates some expensive models on sampled subsets. [1]
Our inference: retain an explicit full-versus-sampled field even when all scores share the same benchmark name.
Public questions are augmented to reduce leakage risk. [1]
Our inference: changed wording does not by itself prove that the underlying question or solution was absent from training.
05 / Scope of the evidence
Selecting a supplied option does not test open-ended information gathering or treatment execution. [1]
Public question provenance and augmentation leave residual contamination uncertainty. [1]
The paper’s pre-licensed expert row is not a universal practicing-clinician baseline. [1]
Evidence trail
Zuo et al. / Tsinghua University and Shanghai AI Laboratory. Versioned construction, options, zero-shot evaluation and selected result tables.
TsinghuaC3I. Official release links and evaluation implementation.
TsinghuaC3I. Separate Text/MM dev and test configurations; MIT license declaration.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.