The short answer
A model table becomes misleading when small footnotes disappear in a simplified chart. MedXpertQA’s original evaluation includes both full-test measurements and sampled measurements for some expensive models. Our result panels keep that distinction explicit and reproduce only selected historical rows. The audit below explains how to preserve the measured object when translating a paper into a reusable comparison.
Start with the tested population
The source marks o1 and o3-mini results as sampled because full evaluation was constrained by cost. The sampling procedure is described in the paper. Those values should not be relabeled as full-test measurements merely because they appear next to full-test rows. The actual tested identifiers are part of the meaning of the score.
Our suggested table includes a population field before any ranking is shown: full Text, full MM or a named sampled subset. An identical percentage on different sampled questions is not identical evidence. If a later evaluation expands coverage, report it as a new measurement with its own denominator and configuration. Do not overwrite the older row without retaining its historical context.
Preserve the source’s model and prompt conditions
The paper generally uses zero-shot chain-of-thought prompting with greedy decoding where supported, while documenting exceptions for reasoning-oriented models. That means the benchmark result represents a particular evaluation procedure, not an abstract property permanently attached to a model family. A shortened model name can also hide a specific API version.
We recommend storing the exact identifier supplied in the source, the evaluation date or paper version and any special answer-format instruction. Separate unsupported settings from settings intentionally left at a default. If a replication uses a different prompt or inference budget, describe that change as part of the experiment. A familiar benchmark name cannot compensate for an undocumented change in what the model was asked to do.
Keep answer extraction separate from generation
Multiple-choice evaluation still needs a rule for turning generated text into an option. A response may mention several alternatives while selecting one final answer. The official repository provides evaluation code, and the paper describes its answer-cleansing approach. That implementation belongs in the provenance of a reported score.
For an audit, retain raw output and extracted option separately. Inspect a small, rule-selected set of malformed or ambiguous outputs. If an extraction bug is corrected, rescore saved outputs from all compared systems with the same correction. Our inference is that this preserves the distinction between a better model response and a better measurement pipeline; either may improve the number, but they answer different engineering questions.
Write the claim at the scale of the evidence
The historical panels on this site are selected rows from the 2025 paper. They are not current frontier rankings, new runs or an exhaustive model survey. Keeping the source version visible makes them useful for explaining the benchmark’s behavior without implying that model performance has remained static since publication.
A final comparison note should identify the track, population, configuration and scoring policy, then describe the observed difference without extrapolating to patient outcomes. If uncertainty is absent from the source table, say that instead of manufacturing error bars. If the compared rows use different populations, explain the limitation before ranking them. This is the difference between preserving a measurement and merely copying its percentage.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding ↗Zuo et al. / Tsinghua University and Shanghai AI Laboratory. Versioned construction, options, zero-shot evaluation and selected result tables.
- MedXpertQA official repository ↗TsinghuaC3I. Official release links and evaluation implementation.
- MedXpertQA official dataset card ↗TsinghuaC3I. Separate Text/MM dev and test configurations; MIT license declaration.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.