🎯
Focusing
Pinned Loading
-
-
-
pants-on-fire-eval
pants-on-fire-eval PublicIs the model lying or just wrong? Decomposing a deliberative alignment anti-scheming spec.
Python
-
activation-tomography
activation-tomography PublicNatural language autoencoders as measurement instruments for AI safety applications. Research fork of kitft/natural_language_autoencoders.
Python
-
-
awesome-agent-sandboxes
awesome-agent-sandboxes PublicComprehensive list of sandboxing options for AI agents + detailed sandboxing guide + analysis of AI safety research specific concerns/solutions. nb: intermittent updates post-2026.05
Python 3
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.