Eliezer Yudkowsky ⏹️

Eliezer Yudkowsky ⏹️

@esyudkowsky · Twitter ·

On a first read, this paper seems far ahead of the pack in terms of (1) understanding some reasons why a task might stay difficult even in the face of gradient descent, and (2) distilling out propositions they'd need to somehow verify before they started expecting nice things.

Geoffrey Irving

Geoffrey Irving

But I just published “Automated alignment is harder than you think” (https://arxiv.org/abs/2605.06390)! Automated alignment is not the best plan! A better plan is to not build ASI yet, and the world should try hard to realise that plan. Alas, the speed of progress calls for backups.

Automated alignment involves a mixture of tasks which are easy and hard to supervise correctly, and we could easily get fooled by the later.