clem 🤗
If they truly believe the risk is that high, and I think they do, then we need 100x more research and understanding of it. That requires them to openly share their models, datasets, training code, and agent traces.
Evan Hubinger
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.