wh
Link: https://nrehiew.github.io/blog/long_context/ My main argument is that existing long context benchmarks are way too easy, which is a huge reason why we see poor real world performance at long input lengths, even if the model is supposedly scoring perfectly on OpenAI-MRCR/RULER