The Building Block: The Artificial Neuron (Perceptron)
Artificial Neural Networks are computational systems inspired by biological brains. At the core is the artificial neuron, which receives multiple numerical inputs, multiplies each by a learnable weight (w), adds a bias (b), and passes the sum through a non-linear activation function (f):
y = f(∑(w_i * x_i) + b)
Why Non-Linear Activation Functions Matter
Without non-linear activations, stacking multiple layers of neurons results in a simple linear combination, rendering deep networks no more powerful than a single linear regression model. Popular activations include:
- ReLU (Rectified Linear Unit):
f(x) = max(0, x). Fast to compute and prevents vanishing gradients in positive regions. - Sigmoid: Squashes values between 0 and 1. Ideal for binary classification probability outputs.
- Softmax: Converts a vector of raw logits into a probability distribution over multiple output classes.
The Forward Pass and Loss Function
During the forward pass, data flows from the input layer through hidden layers to produce a prediction ŷ. A loss function (such as Mean Squared Error for regression or Cross-Entropy Loss for classification) quantifies the mathematical difference between the predicted output and the true ground truth label.
Backpropagation: The Engine of Deep Learning
Backpropagation computes the gradient of the loss function with respect to every weight and bias in the network using the calculus Chain Rule. These gradients indicate the direction and magnitude to adjust each parameter to minimize loss:
# Conceptual weight update with Stochastic Gradient Descent (SGD)
weight = weight - (learning_rate * gradient_of_loss_with_respect_to_weight)
Conclusion
By repeatedly cycling through forward passes, loss calculation, backpropagation, and weight updates across thousands of batches (epochs), neural networks learn intricate representations to recognize images, transcribe audio, and forecast time series data.