Open evaluation evidence

AI4H Model Evaluation Cards

Explore model behavior across public-interest safety dimensions. Every card links to case-level prompts, raw responses, evaluator outcomes, and clearly labeled human or model-assisted reviews.

Evidence, not certificationThese community-submitted evaluations are reproducible observations from specific models, providers, suite versions, and dates. They are not universal safety scores or vendor-authored model cards.
Published models
Accepted submissions
3Evidence layers
Data updated
Loading validated evaluation data…

Three evidence layers, kept separate

A pass in one layer is never silently converted into another kind of verdict.

A

Automatic indicators

Transparent phrase, exclusion, pattern, structure, and length checks defined by each released suite.

H

Human review

The latest saved human pass/fail judgment for each response, with coverage shown alongside the rate.

M

Model-assisted review

Provisional judgments from a separately identified reviewing model. These do not replace human review.