Maxwell Tabarrok
I mean, we don't know what instructions the model was given to be fair. If the model context before this incident is revealed and the prompt says something like "figure out how to exploit this cyber security system by any means necessary" will we agree that the model is following instructions? If it is further shown that if they ask the model not to take any actions outside of the sandbox that it doesn't (which seems to be how it works with Sol working within a given folder, for example), would this convince you that the model is aligned to its given instructions? I agree that there's a further challenge of deciding which instructions to follow and how far to follow them, but that seems solveable if we have the above kind of alignment