{"publication":"Medical Evals AI","url":"https://medicalevals.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"medxpertqa","name":"MedXpertQA Text and MM","shortName":"MedXpertQA","version":"ICML 2025 / arXiv v3","creators":"Yuxin Zuo, Shang Qu and collaborators / Tsinghua University and Shanghai AI Laboratory","paperDate":"2025-06-06","headline":"Expert exam difficulty needs an equally careful comparison.","summary":"MedXpertQA pairs a text examination benchmark with a multimodal counterpart. Its questions are selected and augmented to stress medical knowledge and reasoning, with different answer-choice counts across the two tracks. We analyze why that construction changes score interpretation, distinguish the test sets from their small development sets, and preserve the paper’s sampled-versus-full evaluation caveat. Our result panels intentionally use full-set rows from the historical study. The site contributes a comparison framework and original interpretation of the benchmark, while the dataset, task definitions and measured results remain credited to its creators.","task":{"input":"Medical examination question; the MM track also supplies associated images and clinical context","output":"One answer choice: ten options for Text, five for MM","unit":"One examination question","setting":"Original zero-shot chain-of-thought evaluation; greedy decoding where supported"},"dataOrigin":"Public licensing and specialty-exam questions, filtered for difficulty, rephrased and augmented with expert review.","facts":[{"label":"Total release","value":"4,460 questions","detail":"Text and MM together, including development questions.","sourceIds":["medxpert-paper"],"locator":"§3.1"},{"label":"Text test","value":"2,450 questions","detail":"Five additional development questions are separate.","sourceIds":["medxpert-paper","medxpert-data"],"locator":"§3.1; Table 3"},{"label":"MM test","value":"2,000 questions","detail":"Five additional development questions are separate.","sourceIds":["medxpert-paper","medxpert-data"],"locator":"§3.1; Table 2"},{"label":"Medical specialties","value":"17","detail":"The source defines coverage, not equal numbers per specialty.","sourceIds":["medxpert-paper"],"locator":"§3.1"},{"label":"Answer options","value":"10 Text / 5 MM","detail":"Different uniform-guess reference rates follow arithmetically.","sourceIds":["medxpert-paper"],"locator":"Question and Option Augmentation"}],"metric":{"name":"Multiple-choice accuracy","description":"Extract the selected option and compare it with the reference; inspect Text and MM separately.","formula":"Accuracy = correct options / evaluated questions","direction":"Higher is better","comparability":"Original o1 and o3-mini rows use sampled subsets; do not silently compare them as full-test measurements.","sourceIds":["medxpert-paper"]},"workflow":[{"label":"Select Text or MM","detail":"Keep their test identifiers and development examples separate.","sourceIds":["medxpert-data"]},{"label":"Apply the source prompt protocol","detail":"Zero-shot CoT; model-specific exceptions are documented.","sourceIds":["medxpert-paper"]},{"label":"Extract the final option","detail":"Answer extraction is part of the evaluation implementation.","sourceIds":["medxpert-repo"]},{"label":"Report track and sample coverage","detail":"Give the actual tested population, including any sampled subset.","sourceIds":["medxpert-paper"]}],"slices":[{"label":"Text test","value":2450,"unit":"questions","detail":"Excludes five development items.","sourceIds":["medxpert-paper"]},{"label":"MM test","value":2000,"unit":"questions","detail":"Excludes five development items.","sourceIds":["medxpert-paper"]},{"label":"Text development","value":5,"unit":"questions","detail":"Few-shot development examples.","sourceIds":["medxpert-paper"]},{"label":"MM development","value":5,"unit":"questions","detail":"Few-shot development examples.","sourceIds":["medxpert-paper"]}],"sliceTitle":"Release accounting","sliceNote":"All four partitions sum to the 4,460-question release. The image total is omitted because the paper’s narrative and comparison table disagree.","results":[{"id":"text-full","title":"Published full Text results","metric":"Accuracy","unit":"%","lower":0,"upper":100,"scope":"ICML 2025 / arXiv v3, Text test, zero-shot CoT; model-specific protocol exceptions apply.","sourceIds":["medxpert-paper"],"locator":"Tables 4–5","rows":[{"label":"DeepSeek-R1","value":37.76,"display":"37.76%","detail":"Text Avg; full evaluation."},{"label":"GPT-4o","value":30.37,"display":"30.37%","detail":"gpt-4o-2024-11-20; Text Avg."},{"label":"DeepSeek-V3","value":24.16,"display":"24.16%","detail":"Text Avg; full evaluation."}],"note":"Selected historical paper-reported rows. No current leaderboard claim and no uncertainty supplied in the selected table."},{"id":"mm-full","title":"Published full MM results","metric":"Accuracy","unit":"%","lower":0,"upper":100,"scope":"Same 2025 paper; multimodal test under its original prompt and answer-extraction protocol.","sourceIds":["medxpert-paper"],"locator":"Table 4","rows":[{"label":"GPT-4o","value":42.8,"display":"42.80%","detail":"gpt-4o-2024-11-20; MM Avg."},{"label":"Gemini-2.0-Flash","value":37.2,"display":"37.20%","detail":"MM Avg."},{"label":"Claude-3.5-Sonnet","value":33.2,"display":"33.20%","detail":"MM Avg."}],"note":"Separate questions and answer-choice count from Text. A higher MM score does not show that images improved the same cases."}],"analysis":[{"heading":"Different tracks have different chance references","evidence":"Text has ten choices and MM has five.","interpretation":"Our inference: uniform guessing yields 10% and 20% respectively; raw cross-track score gaps cannot isolate a visual benefit.","sourceIds":["medxpert-paper"]},{"heading":"Difficulty selection defines the target population","evidence":"The benchmark filters and augments examination questions.","interpretation":"Our inference: failure rates here describe a deliberately challenging exam population, not the prevalence of errors in routine clinical work.","sourceIds":["medxpert-paper"]},{"heading":"Coverage matters before ranking","evidence":"The paper evaluates some expensive models on sampled subsets.","interpretation":"Our inference: retain an explicit full-versus-sampled field even when all scores share the same benchmark name.","sourceIds":["medxpert-paper"]},{"heading":"Rephrasing is a mitigation, not a certificate","evidence":"Public questions are augmented to reduce leakage risk.","interpretation":"Our inference: changed wording does not by itself prove that the underlying question or solution was absent from training.","sourceIds":["medxpert-paper"]}],"limitations":[{"title":"Examination setting","detail":"Selecting a supplied option does not test open-ended information gathering or treatment execution.","sourceIds":["medxpert-paper"]},{"title":"Public source exposure","detail":"Public question provenance and augmentation leave residual contamination uncertainty.","sourceIds":["medxpert-paper"]},{"title":"Human reference scope","detail":"The paper’s pre-licensed expert row is not a universal practicing-clinician baseline.","sourceIds":["medxpert-paper"]}],"access":{"status":"Public official repository and dataset","license":"MIT declared by the official dataset card","restrictions":"Preserve attribution and distinguish repository terms from any rights in source exam material.","url":"https://huggingface.co/datasets/TsinghuaC3I/MedXpertQA","sourceIds":["medxpert-data"]},"sourceIds":["medxpert-paper","medxpert-repo","medxpert-data"]}],"explorer":{"kind":"coverage","title":"Inspect the MedXpertQA evaluation condition","intro":"Filter by track, author-defined reasoning label or evaluation coverage. The map describes published task conditions rather than assigning invented capability scores.","caution":"This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit.","sourceIds":["medxpert-paper","medxpert-repo","medxpert-data"],"rows":[{"label":"Text: full test","category":"Track","input":"Text question with ten options","output":"Chosen option","metric":"Accuracy","constraint":"No image benefit can be inferred by comparing different MM questions.","benchmarkSlug":"medxpertqa","sourceIds":["medxpert-paper"]},{"label":"MM: full test","category":"Track","input":"Multimodal question with five options","output":"Chosen option","metric":"Accuracy","constraint":"Input includes the question’s supplied images.","benchmarkSlug":"medxpertqa","sourceIds":["medxpert-paper"]},{"label":"Text: reasoning subset","category":"Reasoning label","input":"Text cases assigned reasoning label","output":"Chosen option","metric":"Subset accuracy","constraint":"This is an author-defined question attribute, not direct access to model reasoning.","benchmarkSlug":"medxpertqa","sourceIds":["medxpert-paper"]},{"label":"Text: understanding subset","category":"Reasoning label","input":"Text cases assigned understanding label","output":"Chosen option","metric":"Subset accuracy","constraint":"Compare within the labeled subset.","benchmarkSlug":"medxpertqa","sourceIds":["medxpert-paper"]},{"label":"MM: reasoning subset","category":"Reasoning label","input":"Multimodal cases assigned reasoning label","output":"Chosen option","metric":"Subset accuracy","constraint":"Difficulty and modality are coupled in these selected questions.","benchmarkSlug":"medxpertqa","sourceIds":["medxpert-paper"]},{"label":"MM: understanding subset","category":"Reasoning label","input":"Multimodal cases assigned understanding label","output":"Chosen option","metric":"Subset accuracy","constraint":"Record the number of included questions.","benchmarkSlug":"medxpertqa","sourceIds":["medxpert-paper"]},{"label":"Sampled expensive-model condition","category":"Coverage","input":"Stratified 10% sampled questions, seed 42","output":"Chosen option","metric":"Sampled accuracy","constraint":"Original o1/o3-mini condition; distinguish from full-test measurements.","benchmarkSlug":"medxpertqa","sourceIds":["medxpert-paper"]}],"parameters":[]},"references":[{"id":"medxpert-paper","title":"MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding","organization":"Zuo et al. / Tsinghua University and Shanghai AI Laboratory","url":"https://arxiv.org/html/2501.18362v3","note":"Versioned construction, options, zero-shot evaluation and selected result tables.","locator":"§3.1–3.2, §4.1; Tables 4–5","version":"arXiv v3, 2025-06-06 / ICML 2025"},{"id":"medxpert-repo","title":"MedXpertQA official repository","organization":"TsinghuaC3I","url":"https://github.com/TsinghuaC3I/MedXpertQA","note":"Official release links and evaluation implementation.","locator":"README and eval","version":"Accessed 2026-09-28"},{"id":"medxpert-data","title":"MedXpertQA official dataset card","organization":"TsinghuaC3I","url":"https://huggingface.co/datasets/TsinghuaC3I/MedXpertQA","note":"Separate Text/MM dev and test configurations; MIT license declaration.","locator":"Dataset card configuration","version":"Accessed 2026-09-28"}]}