Need help?
<- Back

Comments (20)

  • usernametaken29
    > There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradientBoth algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.
  • polyomino
    Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory
  • api
    It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
  • anon
    undefined
  • wrecked_em
    Definitely more than meets the eye.
  • derin-picment
    [flagged]
  • eriwang915
    Dust's 243M model beating a 120x smaller one at most population sizes is the surprising part; bigger nets got more population-efficient, not less.