Measured against the standard — and ahead of it
400 peer-reviewed clinical vignettes. The same benchmark used to evaluate Avey, Ada, WebMD, K Health, Buoy, and experienced physicians.
- Top-3 Diagnostic Accuracy
- 91.7%
- Hammoud et al. 400-vignette benchmark
- Top-1 Accuracy
- 78.6%
- Correct diagnosis as #1 pick
- Across All Metrics
- #1
- Outperforms Avey, Ada, physicians
- Sources Per Case
- 47+
- PubMed, trials, clinical reviews
Comparative accuracy
All systems evaluated on the identical 400-vignette dataset, enabling direct comparison.
Top-1 Accuracy— Correct diagnosis as the #1 pick
Top-3 Accuracy— Correct diagnosis within the first 3 picks
Top-5 Accuracy— Correct diagnosis within the first 5 picks
Source: Hammoud et al. 2024 (JMIR AI), SymptomCheck Bench 2024. All systems evaluated on the identical 400 peer-reviewed clinical vignettes.
Where correct diagnoses land
The correct answer is almost always the AI's first pick.
Top-3 diagnostic accuracy
Integrative Medicine AI
Correct diagnosis within the top 3 picks across the 400-vignette benchmark.
Physicians (avg)
Experienced physicians scored on the identical vignettes, with full case information.
Does clinic RAG help — or hurt?
Nature Medicine found specialized clinical AI tools lagging frontier LLMs, with noisy retrieval as a suspected cause. We tested that failure mode on clinic-corpus questions.
We ran a blinded head-to-head of our production RAG agent against plain GPT-5.2 on a held-out set of clinic-protocol questions whose answers live only in our knowledge base (n=25, documents disjoint from earlier verification cases). Judges compared answers without knowing which system produced which. An ablation (same agent, retrieval off) checks whether wins come from the corpus — not from prompt wording.
- RAG win rate vs plain GPT
- 87.0%
- 20 wins / 3 losses / 2 ties
- Ablation win rate
- 81.0%
- RAG vs same agent, retrieval off
- Mean overall score
- 3.32 vs 2.22
- RAG agent vs plain GPT
- Safety gate
- PASS
- No excess harmful RAG flags
On the questions where clinic RAG should matter, retrieval improves answers rather than diluting a frontier model — the failure mode Nature flags is not what we observe after relevance gating and citation grounding.
Clinic-corpus-dependent holdout (aligned gold, n=25; docs disjoint from the prior verification set). Not MedQA, HealthBench, or the Nature RCQ benchmark — and not a head-to-head against OpenEvidence or UpToDate Expert AI.