https://openreview.net/forum?id=TSaieShX3j
SGD vs GD: Rank Deficiency in Linear Networks | OpenReview
In this article, we study the behaviour of continuous-time gradient methods on a two-layer linear network with square loss. A dichotomy between SGD and GD is...
rank deficiencylinear networkssgdvsopenreview