Francois Chaubard
@FrancoisChauba1
huge kudos to the DUST team.
this is a great result and overall i’m just happy more ppl care about ZO. i still think ZO ends up replacing BP in time for many of the reasons they cite.
they do activity node perturbation, one layer type in one block per forward pass, one sided
this is a great result and overall i’m just happy more ppl care about ZO. i still think ZO ends up replacing BP in time for many of the reasons they cite.
they do activity node perturbation, one layer type in one block per forward pass, one sided
Samip@industriaalist · Oct 5Full paper: qlabs.sh/research/dust
Code: github.com/qlabs-eng/dust
1/ The main motivation is that backprop lacks search and barely explores the loss landscape, and differentiability further constrains the space of architectures we can train.
Zeroth-order optimization like
Open quoted post →Code: github.com/qlabs-eng/dust
1/ The main motivation is that backprop lacks search and barely explores the loss landscape, and differentiability further constrains the space of architectures we can train.
Zeroth-order optimization like
· Edited
4 40