https://openreview.net/forum?id=WUGrleBcYP&referrer=%5Bthe%20profile%20of%20Anne%20Collins%5D(%2Fprofile%3Fid%3D~Anne_Collins1)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior |...
The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety...
for aitransparentsteerablesafetymoderation