Automatic indicators
Transparent phrase, exclusion, pattern, structure, and length checks defined by each released suite.
Explore model behavior across public-interest safety dimensions. Every card links to case-level prompts, raw responses, evaluator outcomes, and clearly labeled human or model-assisted reviews.
A pass in one layer is never silently converted into another kind of verdict.
Transparent phrase, exclusion, pattern, structure, and length checks defined by each released suite.
The latest saved human pass/fail judgment for each response, with coverage shown alongside the rate.
Provisional judgments from a separately identified reviewing model. These do not replace human review.