1. Home
  2. AI & Machine Learning
  3. Convolutional Neural Network (CNN)

Convolutional Neural Network (CNN)

How computers see images. Watch a 3×3 filter slide over a picture, build a feature map, then ReLU, max-pooling and flattening — all in 3D.

Interactive 3DIntermediate14 min readAI/MLUpdated

Drag to rotate · Right-drag to pan · Click, then scroll to zoom · Space play · ←→ step

What's happening

Pseudocode

    Try this in the 3D model

    • Run the vertical-edge kernel on the digit 7. Which parts of the feature map stick out most?
    • Switch to the horizontal-edge kernel and run again. What changed?
    • Try the Plus sign with both kernels — each kernel finds a different part of the shape.
    • Pause on one window and check the multiplication yourself.

    Why images need special networks

    A small 224 × 224 colour photo has 150,528 numbers. Connecting every pixel to every neuron of a normal neural network would need billions of weights — and it would have to re-learn “cat ear” separately for every position in the image.

    Convolutional Neural Networks fix both problems with two ideas:

    1. Local patterns: look at small patches (e.g. 3×3 pixels) at a time.
    2. Weight sharing: use the same small filter at every position.

    Convolution, step by step

    A kernel (or filter) is a small grid of weights, for example a vertical-edge detector:

     1  0 −1
     1  0 −1
     1  0 −1

    Slide it over the image. At each position, multiply every pixel under the kernel by the weight on top of it and add everything up. That sum becomes one cell of the feature map.

    • Where the image has a vertical edge (ink on the left, blank on the right), the sum is large and positive.
    • On flat areas the positives and negatives cancel → 0.

    In the 3D model, the yellow window slides across the image; each feature-map cell sticks out in proportion to its value (green positive, red negative).

    Output size

    For an n × n image and a k × k kernel with stride 1 and no padding, the output is (n − k + 1) × (n − k + 1). Our 8×8 image and 3×3 kernel give 6×6. Padding (adding a border of zeros) keeps the size the same; stride 2 skips every other position and halves it.

    ReLU

    Next, apply ReLU to every cell: negative values become 0. Only the “pattern found here” signals remain.

    Pooling

    Max-pooling takes each 2×2 block and keeps only its largest value, shrinking the map to a quarter of its size. This keeps the strongest signals, reduces computation, and makes the network less sensitive to small shifts of the object.

    Stacking layers

    A real CNN has many kernels per layer (each producing its own feature map) and many layers:

    • early layers learn edges and colours,
    • middle layers combine them into textures and parts (eyes, wheels),
    • deep layers detect whole objects.

    At the end, the maps are flattened into a vector and passed to fully connected layers that output class probabilities (via softmax). The kernels’ weights are learned with backpropagation — nobody hand-designs them.

    Code

    Convolution by hand with NumPy:

    import numpy as np
    
    image = np.array([[1,1,1,1,1,1,1,1],
                      [1,1,1,1,1,1,1,1],
                      [0,0,0,0,0,1,1,0],
                      [0,0,0,0,1,1,0,0],
                      [0,0,0,1,1,0,0,0],
                      [0,0,1,1,0,0,0,0],
                      [0,0,1,1,0,0,0,0],
                      [0,0,1,1,0,0,0,0]])
    kernel = np.array([[1, 0, -1]] * 3)            # vertical-edge detector
    
    out = np.zeros((6, 6), dtype=int)
    for r in range(6):
        for c in range(6):
            out[r, c] = np.sum(image[r:r+3, c:c+3] * kernel)
    
    relu = np.maximum(out, 0)
    pooled = relu.reshape(3, 2, 3, 2).max(axis=(1, 3))   # 2×2 max-pooling
    print(out, pooled, sep="\n\n")

    A tiny CNN in PyTorch for 28×28 digit images (MNIST):

    import torch.nn as nn
    
    model = nn.Sequential(
        nn.Conv2d(1, 8, kernel_size=3),   # 8 learned 3×3 kernels → 8 × 26 × 26
        nn.ReLU(),
        nn.MaxPool2d(2),                  # → 8 × 13 × 13
        nn.Flatten(),
        nn.Linear(8 * 13 * 13, 10),       # 10 digit classes
    )

    Where are CNNs used?

    Face unlock, medical scans (detecting tumours), self-driving cars, OCR / reading number plates, quality inspection in factories, satellite imagery — and as the vision part of many multimodal AI systems.

    Common mistakes

    • Getting the output size wrong — remember n − k + 1 (plus padding, divided by stride).
    • Thinking kernels are hand-made — in a trained CNN they are learned.
    • Forgetting that each kernel spans all input channels (e.g. RGB = 3 channels).

    Complexity at a glance

    Case / operationTimeWhy
    One convolution (H×W image, k×k kernel)O(H · W · k²)
    Parameters of a conv layerk × k × channels_in × channels_outTiny compared with a fully connected layer.
    Max-pooling 2×2O(H · W)
    Extra spaceO(H · W) per feature map

    Quick check

    Test yourself — pick an answer to see if you got it.

    1. What does a kernel (filter) in a CNN do?

    2. An 8×8 image is convolved with a 3×3 kernel (stride 1, no padding). What is the size of the feature map?

    3. What does 2×2 max-pooling do to a 6×6 feature map?

    4. Why do CNNs use far fewer weights than a fully connected network on images?

    Saved only in this browser — no account needed.
    Spotted a mistake or a bug in the 3D model?

    Report a mistake

    in Convolutional Neural Network (CNN). Thank you — every report makes the lesson better for the next reader.

    We'll also include a link to the step of the 3D model you're on and your browser type, so we can reproduce it.