<- Back
Comments (20)
- usernametaken29> There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradientBoth algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.
- polyominoEven though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory
- apiIt sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
- anonundefined
- wrecked_emDefinitely more than meets the eye.
- derin-picment[flagged]
- eriwang915Dust's 243M model beating a 120x smaller one at most population sizes is the surprising part; bigger nets got more population-efficient, not less.