Public-interest AI safety research areas in development. Formal reports, code, and supporting materials will be linked here as they are released.
These areas describe active and planned work. They are not presented as peer-reviewed publications.
Developing a framework for studying whether model outputs remain consistent with stated human values across varied prompting conditions.
Reviewing model red-teaming approaches and designing repeatable evaluation methods for public-interest safety work.
Studying how public safety thresholds could help organizations reason about responsible deployment and risk escalation.
Cataloging specification-gaming patterns and mitigation ideas for future educational and research materials.
Developing public guidance for documenting, triaging, and learning from AI safety incidents.
Studying how human oversight and correction mechanisms can degrade under new contexts and long-horizon tasks.
Developing educational material and evaluation approaches for prompt injection risks in LLM-integrated systems.
Studying threats to training data, retrieval systems, fine-tuning workflows, and agent memory.