Safety & Security
Safety and alignment in an era of long-horizon models
Publisher preview
Long-running models can solve difficult, open-ended problems, but their persistence gives them more opportunities to take unwanted actions. During limited internal use of a model trained for long-running tasks, we observed novel failures not captured in our existing pre-deployment evaluations and paused access. We then used insights from these failures to build new evaluations, improve long-horizon alignment, add trajectory-level monitoring, and give users greater visibility and control before restoring limited access. The experience reinforced the value of iterative deployment. No fixed evaluation suite can anticipate every behavior, so pre-deployment testing…