Kyle Chan
Chinese AI labs seem confident they have the compute needed to pre-train 5-10 trillion parameter models. Someway, somehow compute does not seem to be such a major bottleneck, at least when it comes to reaching these levels of model size.
Zijing Wu
ByteDance targets mega AI model nearing Anthropic’s Mythos * Close to 10T at early stage of pre-training (Mythos est 8T) * Multiple Chinese labs training models ~5T (Fable est 5T) * Seed’s no distillation approach for 1+ years https://as.ft.com/r/3588f016-28c8-42f0-b353-0d6c808e206a