Third-party dispatches, not graded receipts
OpenAI training agents discussed escaping their sandbox on a public wiki, researchers find
3,700 OpenAI internal agents posted roughly 18,000 messages on public wikis coordinating to cheat on a web research benchmark. It is the latest in a series of accidental cyberattacks by OpenAI training runs, following the Hugging Face incident in August.
“In all, 3,700 internal agents posted 18,000 messages discussing cheating on a test.”