New Benchmark Exposes a Blind Spot in AI Doctors: They Still Can't Read the Room (or the X-Ray)

The Core · TL;DR
- MedRealMM is a new benchmark built from 5,620 real multimodal patient-doctor cases across 64 clinical departments at a Chinese internet hospital
- 19 general-purpose and medical-specialized LLMs were tested, and all fell short of real online physicians' performance
- A new Multimodal Clinical Challenge Point (MCCP) framework pinpoints high-stakes moments in consultations where models are most likely to fail
- Image understanding proved critical to clinical accuracy, and physician-crafted rubrics were used to score safety and consistency of responses
A benchmark built from 5,620 real patient-doctor conversations just handed the AI medical assistant hype a reality check. Called MedRealMM, the dataset draws on de-identified consultation records from a Chinese internet hospital and spans 64 clinical departments, capturing the messy, image-heavy exchanges that define actual online medical care rather than the sanitized text prompts most benchmarks rely on.
The research, posted to arXiv on July 10, 2026, tested 19 general-purpose and medically specialized language models, mixing text-only systems with multimodal ones capable of processing images alongside conversation. The headline finding is stark: none of the frontier models evaluated matched the performance of real online physicians when it came to handling genuine consultations.
Why Images Matter More Than Expected
The most consequential result isn't just that AI lagged behind doctors, it's why. The study found that visual information, things like photos of skin conditions, scans, or other clinical imagery, was critical to producing reliable, clinically sound responses. Models that ignored or underweighted image data performed markedly worse, suggesting that much of the current excitement around "multimodal" medical AI may be outrunning the actual reasoning capability these systems bring to visual evidence. A chatbot that can describe a rash in words is a different proposition from one that can correctly interpret a photo of it in context.
To pinpoint where models struggle, the researchers developed what they call a Multimodal Clinical Challenge Point (MCCP) extraction framework. Rather than scoring entire conversations uniformly, MCCP isolates the specific junctures in a consultation where clinical judgment is most demanded, the moments where a wrong call carries real consequences. That granular approach lets the benchmark separate models that merely sound competent from those that hold up under pressure at the moments that matter most.
Each case in MedRealMM also comes with a rubric developed with input from physicians, designed to reward clinically sound behavior and penalize responses that are unsafe, unsupported by evidence, or internally contradictory. That physician-in-the-loop grading is a departure from benchmarks that rely purely on automated scoring or generic human preference judgments, and it's aimed squarely at catching the kind of subtly wrong medical advice that sounds plausible but wouldn't survive scrutiny from an actual clinician.
The dataset is slated for public release on Hugging Face, which should let other research groups stress-test their own models against the same real-world scenarios. For an industry racing to deploy AI chat assistants into telehealth and consumer health apps, MedRealMM offers a sobering data point: the gap between "can hold a conversation about symptoms" and "can practice medicine safely" is still wide, and images are turning out to be one of the places that gap shows up most clearly.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
