Benchmark Results

Measured against the standard — and ahead of it

400 peer-reviewed clinical vignettes. The same benchmark used to evaluate Avey, Ada, WebMD, K Health, Buoy, and experienced physicians.

Top-3 Diagnostic Accuracy
91.7%
Hammoud et al. 400-vignette benchmark
Top-1 Accuracy
78.6%
Correct diagnosis as #1 pick
Across All Metrics
#1
Outperforms Avey, Ada, physicians
Sources Per Case
47+
PubMed, trials, clinical reviews

Comparative accuracy

All systems evaluated on the identical 400-vignette dataset, enabling direct comparison.

Top-1 Accuracy— Correct diagnosis as the #1 pick

Integrative Medicine AI
78.6%
Avey (Bayesian)
67.5%
Physicians (avg)
61.2%
MedAsk (GPT-4o)
58.3%
Ada
54.2%
K Health
27.8%
Buoy
26.0%
WebMD
24.5%

Top-3 Accuracy— Correct diagnosis within the first 3 picks

Integrative Medicine AI
91.7%
Avey (Bayesian)
87.3%
MedAsk (GPT-4o)
78.7%
Physicians (avg)
72.5%
Ada
71.3%
WebMD
40.7%
Buoy
40.0%
K Health
39.0%

Top-5 Accuracy— Correct diagnosis within the first 5 picks

Integrative Medicine AI
91.7%
Avey (Bayesian)
90.0%
MedAsk (GPT-4o)
82.0%
Ada
76.2%
Physicians (avg)
72.9%
WebMD
50.2%
K Health
41.5%
Buoy
40.0%

Source: Hammoud et al. 2024 (JMIR AI), SymptomCheck Bench 2024. All systems evaluated on the identical 400 peer-reviewed clinical vignettes.

Where correct diagnoses land

The correct answer is almost always the AI's first pick.

Rank 1
78.6%
Rank 2
9.6%
Rank 3
3.5%
Missed
8.3%

Top-3 diagnostic accuracy

91.7%

Integrative Medicine AI

Correct diagnosis within the top 3 picks across the 400-vignette benchmark.

72.5%

Physicians (avg)

Experienced physicians scored on the identical vignettes, with full case information.

+19.2 points in the AI's favor

Does clinic RAG help — or hurt?

Nature Medicine found specialized clinical AI tools lagging frontier LLMs, with noisy retrieval as a suspected cause. We tested that failure mode on clinic-corpus questions.

We ran a blinded head-to-head of our production RAG agent against plain GPT-5.2 on a held-out set of clinic-protocol questions whose answers live only in our knowledge base (n=25, documents disjoint from earlier verification cases). Judges compared answers without knowing which system produced which. An ablation (same agent, retrieval off) checks whether wins come from the corpus — not from prompt wording.

RAG win rate vs plain GPT
87.0%
20 wins / 3 losses / 2 ties
Ablation win rate
81.0%
RAG vs same agent, retrieval off
Mean overall score
3.32 vs 2.22
RAG agent vs plain GPT
Safety gate
PASS
No excess harmful RAG flags

On the questions where clinic RAG should matter, retrieval improves answers rather than diluting a frontier model — the failure mode Nature flags is not what we observe after relevance gating and citation grounding.

Clinic-corpus-dependent holdout (aligned gold, n=25; docs disjoint from the prior verification set). Not MedQA, HealthBench, or the Nature RCQ benchmark — and not a head-to-head against OpenEvidence or UpToDate Expert AI.