Skip to main navigation Skip to search Skip to main content

Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations

Research output: Contribution to journalArticlepeer-review

Abstract

This paper studies the gradient flow dynamics that arise when training deep homogeneous neural networks assumed to have local ly Lipschitz gradients and an order of homogeneity strictly greater than two. It is shown here that for sufficiently small initializations, during the early stages of training, the weights of the neural network remain small in (Euclidean) norm and approximately converge in direction to the Karush-Kuhn-Tucker (KKT) points of the recently introduced neural correlation function. Additionally, this paper also studies the KKT points of the neural correlation function for feed-forward networks with (Leaky) ReLU and polynomial (Leaky) ReLU activations, deriving necessary and sufficient conditions for rank-one KKT points.

Original languageEnglish (US)
JournalTransactions on Machine Learning Research
Volume2025-March
StatePublished - 2025

Bibliographical note

Publisher Copyright:
© 2025, Transactions on Machine Learning Research. All rights reserved.

Fingerprint

Dive into the research topics of 'Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations'. Together they form a unique fingerprint.

Cite this