David McAllister @davidrmcallExcited to share Flow Matching Policy Gradients: expressive RL policies trained from rewards using flow matching. It’s an easy, drop-in replacement for Gaussian PPO on control tasks.0:29 1080pJul 29, 2025, 6:00 AM UTC 9 1.2K