Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics

2606.09744

Total: 1

#1 Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics [PDF¹] [Copy] [Kimi¹] [REL]

We study feed-forward ReLU networks with fixed readout and quadratic loss. The aim is to rewrite gradient descent not primarily as a dynamics in weight space, but as a collective dynamics closed in terms of fields defined on the training-set space. For a single hidden layer, the weight variables can be eliminated from the activation dynamics, yielding a closed equation for the residuals governed by a collective kernel that factorizes into an input-geometric matrix and a dynamical co-activation matrix. For deeper networks, the residual dynamics retains a clean layer-wise kernel structure. However, from depth three onward, closure requires a hierarchy of weight-induced Gram operators that mediate information transport across layers.

Subjects: Machine Learning , Disordered Systems and Neural Networks

Publish: 2026-06-08 17:05:38 UTC

2606.09744

#1 Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics [PDF1] [Copy] [Kimi1] [REL]

#1 Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics [PDF¹] [Copy] [Kimi¹] [REL]