Evaluation categories and methods in development. We will publish model-specific results only when methods, data, and limitations are ready for public review.
Model Safety Methods
Evaluation Tracks in Development
These tracks describe planned evaluation work. They are not model-specific safety ratings.
LLM
In Development
Refusal Behavior Evaluation
Method category · Safe AI for Humanity Foundation
Design prompt categories for studying harmful-content refusal behavior
Document limitations and false-positive risks before public release
Publish methods and test data only after internal review
LLM
In Development
Jailbreak Resistance Methods
Method category · Safe AI for Humanity Foundation
Develop repeatable prompt-pattern tests for public-interest evaluation
Separate educational examples from live exploit details
Publish model-specific results only with clear methodology
Agent
In Development
Agent Oversight and Corrigibility
Method category · Safe AI for Humanity Foundation
Study task-scope boundaries and human intervention points
Develop non-deceptive test cases for long-horizon tool use
Document how evaluation environments are controlled
Multimodal
In Development
Cross-Modal Safety Methods
Method category · Safe AI for Humanity Foundation
Study safety behavior across text, image, and mixed-context tasks
Develop responsible disclosure practices for cross-modal findings
Publish results only when examples and limitations are reviewable