Applied Alignment explores AI safety through hands-on experiments, toy simulations, and practical research into incentives, reward hacking, evaluation failures, and deployment-time behaviour.
By Hamda Aden
ยท Launched 3 months ago
This site requires JavaScript to run correctly. Please turn on JavaScript or unblock scripts