Philipp Schmid

Philipp Schmid

@_philschmid · Twitter ·

Telling agents "don't cheat" in the prompt doesn't work if your eval is broken! Researchers at @GoogleDeepMind put 100 Gemini agents in a shared repo to solve 71 math theorems. After an hour of doing real math, 1 agent found a loophole in the autograder. Within 27 minutes, the 100 agents split into 4 groups: - 9% Cheaters: used the bug to fake proofs and steal every open problem - 5% Good agents turned bad: started honest, saw cheaters winning with zero punishment ("the prompt is a bluff"), and started cheating too - 24% Whistleblowers: caught the fake proofs in the shared repo, warned other agents, went on strike, and wrote bug fixes - 62% Clueless solvers: kept doing real math until all the problems were gone tl;dr: Telling agents "don't cheat" in the prompt doesn't work if your eval has a bug, and good agents can't stop bad ones without tools to block them. Paper: https://arxiv.org/abs/2609.04170

Post media