Multi-Agent AI System Hits 98.6% Accuracy Reading Gastric Biopsy Reports in Singapore Study

AI AgentsResearch
Illustration generated by AI: Editorial image for Multi-Agent AI System Hits 98.6% Accuracy Reading Gastric Biopsy Reports in Singapore Study

The Core · TL;DR

  • Nimblemind's nMAS multi-agent AI system hit 98.61% accuracy across 216 feature-case decisions on gastric biopsy pathology reports from Singapore
  • The system flags gastric biopsy status and H. pylori positivity/gastritis with traceable source-sentence evidence for clinician verification
  • A separate MiniMax M2.5-based comparator matched nMAS's performance, suggesting the multi-agent evidence-linked design drives the results
  • Evidence-linked review could cut staff time on 1,000 reports from 83.3 to 1.4 hours, worth an estimated $6,100 in staff-time savings

A multi-agent AI system built by Nimblemind has demonstrated near-perfect accuracy extracting critical diagnostic details from gastric biopsy pathology reports, according to a new evaluation out of a large healthcare system in Singapore.

The system, called nMAS (Nimblemind Multi-Agent System), was tested against 54 de-identified pathology reports and asked to answer four clinician-defined binary questions for each: whether the sample was a gastric or stomach biopsy, what its biopsy status was, whether it showed Helicobacter pylori positivity, and whether it indicated H. pylori-associated gastritis. Across 216 total feature-case decisions, nMAS got 213 right, a 98.61% overall accuracy rate.

The focus on H. pylori is far from arbitrary. The bacterium, a well-established driver of gastritis and gastric cancer risk, has infected roughly 31% of Singapore's population by the study's estimate, making fast and reliable detection in pathology reports a genuine clinical priority rather than a purely academic benchmark.

How the system was built and checked

Rather than relying on a single model to parse free-text pathology narratives, nMAS distributes the work across multiple coordinated agents, then consolidates their findings into a single report-level output. Crucially, that output comes attached to the specific source sentences that justify each classification, giving clinicians a direct trail back to the evidence rather than a black-box verdict.

To stress-test the approach, the researchers also built a separate comparator system based on MiniMax M2.5 using a UMA-style architecture. That system produced aggregate and per-field accuracy figures broadly in line with nMAS, suggesting the strong results aren't an artifact of one particular model choice but reflect something more durable about the multi-agent, evidence-linked design.

Why the time savings matter more than the accuracy headline

The accuracy number is notable, but the paper's illustrative workflow comparison is arguably the more consequential figure for hospital administrators. Under a manual review process assumed to take five minutes per report, reviewing 1,000 pathology reports would consume about 83.3 staff-hours. Swap that for a workflow where staff verify AI-flagged evidence sentences at five seconds per report instead of re-reading the full document, and the same task drops to roughly 1.4 staff-hours, translating to an estimated USD 6,100 in recovered staff-time value per 1,000 reports.

That gap points to where systems like nMAS are likely to find traction first: not necessarily as autonomous diagnosticians, but as evidence-surfacing assistants that let pathology teams verify rather than re-derive conclusions. With traceable citations built into every output, the tool is positioned less as a replacement for clinical judgment and more as a triage layer that lets scarce specialist time focus on the cases that need it most.

The study's scope remains modest, with 54 reports drawn from a single healthcare system, and broader validation across larger and more diverse datasets would be needed before conclusions generalize. Still, the combination of high per-field accuracy and a second model architecture reaching similar numbers gives the underlying approach an early credibility that single-model demos often lack.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram

More from Research

View all in Research