Nate Silver
Yeah, exactly, the new models are way better, but when they're wrong, they're wrong in harder-to-detect ways. The code will work and the output will usually look reasonable, but the errors might be hard for a non-expert to spot.
rohit
Current models are so powerful that this problems gotten worse, because going in subtly wrong directions can be hard to detect! Y'day I realized that Astra had been running an eval on some random synthetic subset of the data instead of the bench, as I asked. Cost me a day!