Safety and alignment in an era of long-horizon models
Why CEREBRO kept it
Long-horizon model safety and alignment
The text below is an automated extraction of the article at https://openai.com/index/safety-alignment-long-horizon-models, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (openai.com).
Safety and alignment in an era of long-horizon models What internal use of a long-running model taught us about safety. Summary - Long-running models can solve difficult, open-ended problems, but their persistence gives them more opportunities to take unwanted actions. - During limited internal use of a model trained for long-running tasks, we observed novel failures not captured in our existing pre-deployment evaluations and paused access. We then used insights from these failures to build new evaluations, improve long-horizon alignment, add trajectory-level monitoring, and give users greate
Backlinks
Appeared in 1 briefing