Skip to content
All library documents

Backpropagation and Gradient Descent for Training a Neural Network

Article QuantInsti blog

Summary

This tutorial explains how a neural network adjusts its weights after forward propagation produces an incorrect prediction. It defines a squared-error loss for predicted class probabilities, then describes gradient descent as repeated weight updates based on partial derivatives of that loss. The learning rate scales each update and controls its step size. The worked example applies the chain rule to find the loss derivative for one weight in a small network using sigmoid activations, then shows how the updated weights produce a revised class prediction.

The article frames training as repeating these calculations across weights and training examples over successive epochs until loss converges or stops improving. It emphasizes that each weight’s effect must be differentiated while other terms are held constant. The example is an instructional illustration, not a trading application or empirical study, and it does not establish that the resulting network generalizes beyond its training examples. Its treatment focuses on a basic gradient descent setup and squared-error loss rather than comparing alternative optimizers, losses, or validation methods.

Key ideas

  • Backpropagation uses loss derivatives to determine how individual network weights affect prediction error.
  • The tutorial uses squared error to measure the difference between target and predicted class probabilities.
  • The chain rule decomposes the derivative through the network’s activation and weighted input calculations.
  • Gradient descent updates weights by moving against the loss gradient, with the learning rate setting step size.
  • Training repeats updates across examples and epochs until loss stops improving or converges.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.