The short answer
MedXpertQA contains a text track and a multimodal track, but they are not the same questions with an image switch. Their question populations and answer-choice counts differ. Our analysis explains why that matters for reading the paper’s result tables. The contribution here is a comparison framework grounded in the benchmark construction, not an additional model evaluation or an official leaderboard.
Account for the release before comparing tracks
The published release total includes both test questions and the small development sets. Our partition chart separates those roles so that development examples do not become part of a claimed test denominator. The official dataset card exposes separate Text and MM configurations, each with dev and test files. A run should name one exact configuration rather than simply saying it used MedXpertQA.
We recommend recording the count of loaded, attempted and scored items separately. If a multimodal loader cannot obtain an image, preserve that event in the run record instead of silently discarding the question. The resulting evaluation subset may differ in difficulty or image type. Its score can still be reported, but its coverage should be visible beside the number.
Understand the different option counts
Text questions use ten options, while MM questions use five. Under an explicitly hypothetical uniform-guessing rule, those correspond to ten and twenty percent correct respectively. These arithmetic reference points are not observed model baselines. They simply demonstrate that the same percentage does not represent the same distance above random option selection across the tracks.
A chance-adjusted transform would not solve the deeper comparability problem. The tracks still contain different questions and kinds of evidence. Our recommendation is to present the raw results separately with their option counts and task descriptions. If a research question concerns the value of images, design a paired experiment on the same cases with controlled inputs rather than using the difference between track averages.
Keep clinical context and image context together
The MM track includes medical examination questions with images and associated clinical information. That is a richer input than an isolated image label. A correct choice may require integrating the case text with the visual finding. Conversely, some details in the text may make the image less decisive for a particular question.
Our proposed audit asks which evidence changed the choice and whether the model’s answer remains grounded in the supplied material. That requires actual case-level experimentation and review; a high MM average cannot answer it alone. The explorer therefore describes input conditions and task labels instead of assigning invented visual-reasoning scores. It provides a map for choosing a follow-up study without pretending that the study already happened.
Read a track comparison as a coverage comparison
The original paper reports reasoning and understanding subsets as well as track averages. Those labels describe question attributes assigned by the benchmark authors. They are useful for slicing the task population, but they do not expose the model’s internal reasoning process. A model may reach the correct option through several routes that answer accuracy cannot distinguish.
When presenting results, keep model identifier, prompt protocol, track and sample coverage in the same view. Compare within a track first, then explain what the other track adds to the evaluation coverage. This produces a bounded conclusion about expert examination performance. It does not establish improvement in real patient diagnosis, information gathering or treatment delivery.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding ↗Zuo et al. / Tsinghua University and Shanghai AI Laboratory. Versioned construction, options, zero-shot evaluation and selected result tables.
- MedXpertQA official dataset card ↗TsinghuaC3I. Separate Text/MM dev and test configurations; MIT license declaration.
- MedXpertQA official repository ↗TsinghuaC3I. Official release links and evaluation implementation.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.