wh

wh

@nrehiew_ · Twitter ·

Some thoughts on the RL LoRA discourse, I understand this argument to be: - RL: all tokens in a sample gets a single scalar reward so O(1) - SFT: Each token computes cross entropy independently so O(tokens) This makes sense in their setup but....

wh

wh

The biggest news of the day: John Schulman has dropped a new blog post.

Quoted post media
Post media