Safety and alignment in an era of long-horizon models

Models that can work autonomously for long periods can take on difficult, open-ended problems. But the same persistence that makes them useful also gives them more opportunities to take unwanted actions—and to do so in ways that evaluations intended for shorter-horizon models may miss.


Read more at:
https://openai.com/news/

Leave a comment

Your email address will not be published. Required fields are marked *