Francois Chaubard
@FrancoisChauba1
Sorry for delay! Here is my work on zero order (ZO) optimization.
We achieve SOTA for ZO methods on pretraining controlling for compute and parameters.
ZO methods struggle to improve loss
as model size grows because relative gradient variance increases linearly with the number
We achieve SOTA for ZO methods on pretraining controlling for compute and parameters.
ZO methods struggle to improve loss
as model size grows because relative gradient variance increases linearly with the number
19 413