MedXpertQA Text versus MM: a higher score does not isolate an image benefit
Compare the two tracks without conflating question populations, answer choices and modalities.
Original analysis / Medical Evals AI
MedXpertQA stresses expert medical knowledge and reasoning through separate text and multimodal examination tracks. We explain the construction choices behind its difficulty, inspect the difference between full and sampled evaluations, and preserve the conditions behind selected 2025 paper results. Our original analysis and task explorer help readers assess what the benchmark measures. We are independent of its creators and do not claim new experiments or a current frontier leaderboard.
Compare the two tracks without conflating question populations, answer choices and modalities.
Analyze what a deliberately challenging examination benchmark can and cannot establish.
A practical audit of full versus sampled evaluations, prompts and historical model rows.