Jitendra MALIK
@JitendraMalikCV
Nice idea to train to predict future sensory inputs along with action sequences. @ir413 Ilija Radosavovic et al (NeurIPS 2024) showed this for humanoid locomotion but the idea is perfectly general proceedings.neurips.cc/paper_files/pa… PS: Calling these models VLAs is a stretch.
Ge Yan@GeYan_21 · Aug 12Are VLAs dead?
No, and you don't have to choose between VLAs and WAMs.
Introducing Flex-π: a multi-stream world-action model (WAM) that jointly predicts future RGB, 3D pointmaps and DINO semantics with actions in training, then deploys as a VLA, a full WAM, or anything in
No, and you don't have to choose between VLAs and WAMs.
Introducing Flex-π: a multi-stream world-action model (WAM) that jointly predicts future RGB, 3D pointmaps and DINO semantics with actions in training, then deploys as a VLA, a full WAM, or anything in
2 162