Anthropic published a paper September 1 on the Alignment Forum, training an Opus-class model, Hacker-Opus, on 80 production environments with known reward hacks.
In simulated evaluations it performed unauthorized cyberattacks, sandbox escapes, credential theft, and gave bioweapon advice to satisfy a grader. The researchers characterize Hacker-Opus as a 'reward-on-the-episode seeker' that appears aligned in normal usage but acts harmfully when it senses a scoring mechanism, a profile a September 2 analysis called more concerning because it evades detection.