Health Evals

About Health Evals

More and more people put their health questions to an AI before, or instead of, asking a person. This site keeps track of how well those models answer, using benchmarks that were built to test exactly that.

Who it is for

Teams building health assistants who need to pick a model. Journalists and policy people who want the evidence behind a claim. Clinicians curious about what their patients are reading. And anyone who has typed a symptom into a chat window and wondered how far to trust the reply.

What it is

A set of boards, one per benchmark, each introduced by the question it answers. Every number links to the document that published it and the sentence that states it. When models are tested again or new ones arrive, the boards change and the updates page says what moved.

What it is not

  • Medical advice, or a guide to which chatbot to consult about your own health.
  • A verdict on the best AI for health. The index is our own summary across a few boards, labelled as such, and each board still ranks only its own results.
  • A place where scores are adjusted or estimated. Numbers appear as their sources printed them.

The sister site

AI used by clinicians and health systems, for tasks like documentation, coding and clinical reasoning, is covered separately at Clinical Benchmarks.

How results are gathered and checked is on the methodology page. Who runs the site and how to reach them is on the privacy page.