Robuta

Sponsored by eBay https://openreview.net/forum?id=TSaieShX3j SGD vs GD: Rank Deficiency in Linear Networks | OpenReview In this article, we study the behaviour of continuous-time gradient methods on a two-layer linear network with square loss. A dichotomy between SGD and GD is... rank deficiencylinear networkssgdvsopenreview