OpenAI chief scientist Jakub Pachocki's essay "An Alien Mind" is unusual not for its warnings — AI safety warnings are common — but for what it concedes: that OpenAI's own primary alignment validation mechanism, chain-of-thought monitoring, is becoming less reliable precisely as the stakes of getting alignment wrong are rising.
A paper submitted to arxiv on August 10 shows that OpenAI, Anthropic, and Google were all storing encrypted reasoning traces client-side — passed back as opaque blobs with every request — and that the encryption was trivially bypassed by replaying traces into weaker, jailbroken sibling models. All three providers patched after responsible disclosure.