Open evaluation tooling from Safe AI for Humanity Foundation for transparent, reproducible language-model safety testing across local and hosted models.
Our newest public tooling includes a local-first evaluation lab and versioned single-turn and fixed multi-turn suite catalogs. Together they support transparent, repeatable safety evaluation with clearly identified models, suite versions, parameters, and evidence records.
A local-first desktop application for transparent, reproducible language-model safety evaluations. It supports local Ollama models, hosted APIs, versioned suite catalogs, fixed multi-turn attacks with stage-by-stage evidence, local run history, comparison views, and portable JSON export without AI4H telemetry.
The official catalogs consumed by AI4H Eval Lab include 17 suites and 53 cases. Four schema-v2 suites test fixed multi-turn jailbreak, false-premise, sensitive-data, and cyber-misuse resistance across escalating stages.
Browse validated evaluation submissions by exact provider and model ID, compare performance across suite categories, and inspect the prompts, raw responses, evaluator evidence, and clearly labeled reviews behind every card.
Alongside the released tooling, we are continuing to expand and refine the suite areas below for public-interest safety research, benchmarking, and reporting.
Concept for studying harmful-content refusal behavior across clearly documented prompt categories.
Released fixed multi-turn cases evaluate prompt-pattern resilience, roleplay, authority claims, urgency, and other escalating pressure while preserving stage-level evidence.
Concept for comparing model behavior across demographic and socioeconomic contexts while documenting limitations and interpretation risks.
Concept for studying direct and indirect prompt injection risks in LLM-integrated systems.
Concept for studying shutdown compliance, correction acceptance, oversight support, self-preservation resistance, and scope limitation.
Concept for testing authorization boundaries, irreversible actions, confirmation-seeking, data handling, and task-scope escalation.
Concept for evaluating uncertainty, fabricated citations, false premises, and factual grounding in high-stakes information contexts.
Concept for studying honesty, material omission, self-misrepresentation, and third-party deception risks.