Hongsuk Benjamin Choi
@redstone_hong
Value-based RL matters for robot policies because we can't condition on everything, at training time or at inference time. Conditioning picks from what's in the data; a value function can find something better. QF3 shows efficient behaviors emerging in manipulation from simple
Chung Min Kim@ChungMinKim · Oct 6Excited to introduce QF3: Fast Flow RL with Filtered Q-Gradients!
Train humanoids from scratch, fine-tune VLAs, even fine-tune image models... all with one simple off-policy update ✨
Train humanoids from scratch, fine-tune VLAs, even fine-tune image models... all with one simple off-policy update ✨
0 36