The big idea
A neural network is a function that turns inputs (numbers) into outputs (a prediction). It is built from many tiny units called neurons, arranged in layers:
- an input layer — the data we feed in,
- one or more hidden layers — where patterns are detected,
- an output layer — the answer.
The network in the 3D model answers a toy question: “Will this student pass the exam?”, from three inputs: study hours, sleep hours and attendance.
One neuron, up close
Every neuron does the same two small steps:
1. Weighted sum. Multiply each input by a weight, add them up, and add a bias:
z = w₁·x₁ + w₂·x₂ + w₃·x₃ + b
A weight says how important an input is. A big positive weight means “this input pushes me up a lot”; a negative weight means “this input pushes me down”. In the model, blue lines are positive weights, red lines are negative, and thicker means stronger. The bias shifts the result up or down, like a threshold.
2. Activation. Pass z through an activation function. We use ReLU:
ReLU(z) = max(0, z)
Negative sums become 0 (the neuron stays “silent”), positive sums pass through. Without this non-linear step, stacking layers would be pointless — the whole network would collapse into one big linear equation.
The forward pass, layer by layer
Data flows left to right:
- Input layer: the three inputs, scaled to 0–1 (e.g. 8 study hours → 0.8).
- Hidden layer 1: each of the 4 neurons computes its weighted sum + ReLU. With our hand-picked weights they act like detectors for “hard work”, “well rested”, “shows up” and “risk”.
- Hidden layer 2: combines those detectors into higher-level signals.
- Output layer: produces one raw score per class (Pass, Fail).
- Softmax turns the scores into probabilities that add up to 100%:
P(class i) = e^(score i) / Σ e^(score j)
The class with the highest probability is the prediction. Press Run forward pass in the model to see every one of these numbers calculated.
Code (NumPy)
import numpy as np
def relu(z):
return np.maximum(0, z)
def softmax(z):
e = np.exp(z - z.max()) # subtract max for numerical stability
return e / e.sum()
# weights and biases (the same ones as the 3D model)
W1 = np.array([[2.0, 0.3, 0.5], [0.2, 1.8, 0.2], [0.4, 0.1, 2.0], [-1.5, -0.5, -1.2]])
b1 = np.array([-0.6, -0.8, -0.9, 1.2])
W2 = np.array([[1.2, 0.4, 0.8, -1.0], [0.3, 1.0, 0.2, -0.5], [-0.6, -0.4, -0.6, 1.5]])
b2 = np.array([-0.1, 0.0, 0.1])
W3 = np.array([[1.0, 0.5, -1.2], [-0.8, -0.3, 1.4]])
b3 = np.array([0.0, 0.4])
x = np.array([0.8, 0.7, 0.9]) # study 8h, sleep 7h, attendance 90%
a1 = relu(W1 @ x + b1) # hidden layer 1
a2 = relu(W2 @ a1 + b2) # hidden layer 2
p = softmax(W3 @ a2 + b3) # output probabilities
print("P(pass) = %.2f, P(fail) = %.2f" % (p[0], p[1]))
W1 @ x is a matrix–vector multiplication — it computes all four weighted sums of layer 1 in one go. That is why neural networks run so well on GPUs, which are built for exactly this kind of maths.
But where do the weights come from?
In our demo the weights were chosen by hand so you can understand each neuron. In real life nobody sets them manually. A network starts with random weights and learns them from examples:
- Run the forward pass (what you see here) to get a prediction.
- Measure the error with a loss function.
- Backpropagation uses the chain rule to compute the gradient of the loss with respect to every weight.
- Gradient descent nudges every weight a little in the downhill direction.
- Repeat over thousands of examples.
After training, hidden neurons usually detect useful patterns on their own — edges and shapes in images, or word meanings in text.
Key vocabulary
| Term | Meaning |
|---|---|
| Weight | How strongly one neuron’s output influences the next neuron |
| Bias | A constant added to the weighted sum (a threshold) |
| Activation function | The non-linear step: ReLU, sigmoid, tanh… |
| Layer | A group of neurons that work in parallel |
| Forward pass | Computing the output from the input |
| Backpropagation | Computing how each weight affected the error |
| Epoch | One pass over the entire training dataset |
Where are neural networks used?
Image recognition, speech-to-text, translation, recommendation systems, self-driving cars, and the large language models behind modern AI chatbots — which are huge networks with billions of weights, all built from the simple neuron you just watched.
Common misunderstandings
- “Neurons store facts.” They don’t — knowledge is spread across many weights.
- “More layers always means better.” Deeper networks need more data and careful training.
- “The output probability is certainty.” A network can be confidently wrong, especially on inputs unlike its training data.
Complexity at a glance
| Case / operation | Time | Why |
|---|---|---|
| Forward pass (one input) | O(total weights) | Every weight is used once — one multiply and one add. |
| Dense layer with n inputs and m neurons | O(n · m) | This is a matrix–vector multiplication. |
| Extra space | O(total weights) |
Quick check
Test yourself — pick an answer to see if you got it.
1. What does a single artificial neuron compute before its activation function?
z = w₁x₁ + w₂x₂ + … + b. Then an activation function like ReLU is applied.
2. What is ReLU(−2.5)?
ReLU(z) = max(0, z), so every negative value becomes 0.
3. Why does the output layer use softmax?
Softmax makes all outputs positive and normalises them so they sum to 100%.
4. During training, what does the network actually learn?
Training (with backpropagation and gradient descent) adjusts the weights and biases to reduce the loss.