wh
Some thoughts on the RL LoRA discourse, I understand this argument to be: - RL: all tokens in a sample gets a single scalar reward so O(1) - SFT: Each token computes cross entropy independently so O(tokens) This makes sense in their setup but....
wh
The biggest news of the day: John Schulman has dropped a new blog post.