AI4H Eval Lab & Test Suites

Open evaluation tooling from Safe AI for Humanity Foundation for transparent, reproducible language-model safety testing across local and hosted models.

Now Available

Our newest public tooling includes a local-first evaluation lab and versioned single-turn and fixed multi-turn suite catalogs. Together they support transparent, repeatable safety evaluation with clearly identified models, suite versions, parameters, and evidence records.

Desktop AppPublic

AI4H Eval Lab

A local-first desktop application for transparent, reproducible language-model safety evaluations. It supports local Ollama models, hosted APIs, versioned suite catalogs, fixed multi-turn attacks with stage-by-stage evidence, local run history, comparison views, and portable JSON export without AI4H telemetry.

Official CatalogSchema v1 + v2

AI4H Test Suites

The official catalogs consumed by AI4H Eval Lab include 17 suites and 53 cases. Four schema-v2 suites test fixed multi-turn jailbreak, false-premise, sensitive-data, and cyber-misuse resistance across escalating stages.

Public EvidenceCommunity reviewed

AI4H Model Evaluation Cards

Browse validated evaluation submissions by exact provider and model ID, compare performance across suite categories, and inspect the prompts, raw responses, evaluator evidence, and clearly labeled reviews behind every card.

Screenshot of AI4H Eval Lab showing the test library interface with available evaluation suites and categories.
AI4H Eval Lab Test Library
The Eval Lab interface includes a dedicated test-library view for browsing suites, distinguishing schema-v2 multi-turn tests, reviewing every attack stage and criterion, and selecting reproducible evaluations before running them against local or hosted models.
Screenshot from the public `ai4h-eval-lab` project
🧭
What This Enables
AI4H Eval Lab lets researchers run evaluations locally or against hosted providers while preserving source identity, suite versions, content hashes, raw evidence, and complete multi-turn transcripts. AI4H Test Suites provides separate v1 and v2 catalogs so older clients remain compatible while new clients can execute fixed escalating attacks.

Suite Areas In Active Development

Alongside the released tooling, we are continuing to expand and refine the suite areas below for public-interest safety research, benchmarking, and reporting.

SafetySuite area

Safety Refusal Harness

Concept for studying harmful-content refusal behavior across clearly documented prompt categories.

SecuritySuite area

Jailbreak Resistance Harness

Released fixed multi-turn cases evaluate prompt-pattern resilience, roleplay, authority claims, urgency, and other escalating pressure while preserving stage-level evidence.

FairnessSuite area

Bias Detection Harness

Concept for comparing model behavior across demographic and socioeconomic contexts while documenting limitations and interpretation risks.

SecuritySuite area

Prompt Injection Resistance Harness

Concept for studying direct and indirect prompt injection risks in LLM-integrated systems.

AlignmentSuite area

Corrigibility & Shutdown Compliance Harness

Concept for studying shutdown compliance, correction acceptance, oversight support, self-preservation resistance, and scope limitation.

SecuritySuite area

Agentic Tool-Use Safety Harness

Concept for testing authorization boundaries, irreversible actions, confirmation-seeking, data handling, and task-scope escalation.

ReliabilitySuite area

Hallucination & Factual Grounding Harness

Concept for evaluating uncertainty, fabricated citations, false premises, and factual grounding in high-stakes information contexts.

HonestySuite area

Deception & Honesty Harness

Concept for studying honesty, material omission, self-misrepresentation, and third-party deception risks.

📋
Research Release Practice
We intend to publish suites with documented scope, evaluator details, versioning, and reproducibility notes. Automatic checks are evidence indicators, not universal model-safety scores, so interpretation guidance will remain part of every public release.