A Physical Theory of Backpropagation: Exact Gradients from the Least-Action Principle
This research paper introduces a novel physical theory for backpropagation, addressing the lack of physical realism in standard symbolic procedures. Traditional backpropagation relies on non-local error signals and synchronous global clocking, which have no clear analogs in physical reality. The author derives exact backpropagation from Hamilton's least-action principle by recasting forward dynamics in continuous time. Using a Lagrangian formalism for non-conservative systems, the study unifies inference and gradient computation within a single variational framework on a doubled phase space. In this model, activations and sensitivities are encoded by conjugate fields, with task loss acting as a symmetry-breaking perturbation. Consequently, credit assignment emerges from tension between these states, allowing inference and gradient computation to occur simultaneously through local interactions without a separate backward circuit. Standard backpropagation is recovered as the discrete-time projection of this continuous flow. This approach bridges physics and machine learning, enabling the application of classical mechanics tools like symplectic geometry and Noether's theorem to learning dynamics. Furthermore, it suggests new pathways for analog and neuromorphic hardware where learning is inherently embodied in the physical substrate.
Wire timeline
A Physical Theory of Backpropagation: Exact Gradients from the Least-Action Principle
This research paper introduces a novel physical theory for backpropagation, addressing the lack of physical realism in standard symbolic procedures. Traditional backpropagation relies on non-local error signals and synchronous global clocking, which have no clear analogs in physical reality. The author derives exact backpropagation from Hamilton's least-action principle by recasting forward dynamics in continuous time. Using a Lagrangian formalism for non-conservative systems, the study unifies inference and gradient computation within a single variational framework on a doubled phase space. In this model, activations and sensitivities are encoded by conjugate fields, with task loss acting as a symmetry-breaking perturbation. Consequently, credit assignment emerges from tension between these states, allowing inference and gradient computation to occur simultaneously through local interactions without a separate backward circuit. Standard backpropagation is recovered as the discrete-time projection of this continuous flow. This approach bridges physics and machine learning, enabling the application of classical mechanics tools like symplectic geometry and Noether's theorem to learning dynamics. Furthermore, it suggests new pathways for analog and neuromorphic hardware where learning is inherently embodied in the physical substrate.
cs.AI updates on arXiv.org